arXiv ScienceSearch

arXiv subjects

Di Liu

Publications and source records attributed to Di Liu.

At least 19 recordsLinked to original sources

Phonon-assisted transport and hole-phonon coupling in GaAs double quantum dots

Hole-phonon interactions play an important role in transport and decoherence processes in semiconductor quantum dots. Here we investigate hole-phonon coupling in a gate-defined GaAs double quantum dot integrated with a quantum point contact charge sensor. Under finite source-drain bias, pronounced oscillatory stripe patterns appear near specific charge transition regions in the charge stability diagram. We attribute these oscillations to phonon emission during inelastic interdot tunneling. A theoretical model including piezoelectric hole-phonon coupling reproduces the observed patterns. Furthermore, our analysis shows that the oscillations emerge only in particular charge configurations. Our results provide direct insight into phonon-assisted transport and coherent hole-phonon interactions in semiconductor quantum dots.

cond-mat.mes-hall

An Adaptive Longitudinal Platooning Design Based On Concurrent Learning

This work proposes a new adaptive longitudinal platooning strategy in the framework of concurrent learning. Adaptive refers to vehicles facing uncertainty in powertrain parameters via on-line estimation; concurrent learning refers to using both current and past data in the estimation. The proposed platooning strategy advances existing ones since convergence to the true powertrain parameters is guaranteed without imposing persistence of excitation on the vehicle behavior: it suffices the presence of a single non-zero data sample. Meanwhile, the concurrent learning proof we give advances existing ones since it takes into account an extra unknown gain in the error dynamics.

eess.SY

Correct Online Estimation of the Powertrain Time Constants in Adaptive Vehicular Platooning

In longitudinal platooning, some key sources of uncertainty are the powertrain time constants of the vehicles. Because such time constants appear in the input matrix of the platooning dynamics, their correct estimation is either impractical with methods requiring persistence of excitation, or impossible with methods requiring the input matrix to be known. This work proposes a novel adaptive longitudinal platooning method with correct estimation of the powertrain time constants. To achieve correct estimation, the composite adaptive control framework and its stability analysis are suitably modified to handle the time constant uncertainty in the design of the adaptive law. The result is a platooning protocol that guarantees convergence of the estimated time constants to their true values without the need for persistence of excitation: it is sufficient the derivative of the acceleration to be nonzero over a possibly short transient, an extremely relaxed excitation condition. Comparisons with state-of-the-art platooning solutions reveal advantages such as no required measurements of acceleration derivative nor collection of past data. The robustness and practicality of the proposed design is also verified with CarSim-based platooning experiments.

eess.SY

On-chip generation of multi-qubit graph states with high-dimensional encoded single photons

Photonic multi-qubit entanglement is key to optical quantum information processing, particularly universal quantum computing. Yet multi-photon sources suffer from low emission efficiency, making single-photon high-dimensional encoding an appealing alternative. Here we propose an explicit and resource-efficient high-dimensional encoding approach to achieve the target multi-qubit quantum state. The technically challenging preparation of multi-photon quantum states is replaced by single-photon operations involving high-dimensional expansion, routing, and multi-layered quantum measurement. Besides, each photon in the resource multi-photon quantum state can be used to encode multiple qubits in a distributed manner, and a larger entangled state will be constructed. We demonstrate this approach using programmable photonic integrated circuits, where multi-qubit graph states--including the Greenberger-Horne-Zeilinger state and the cluster state--are generated and characterized. We additionally demonstrate the Grover search algorithm using the single-photon cluster state. Our findings unlock a novel route towards diverse entangled state generation with photons and advance large-scale and universal photonic quantum information processing.

quant-ph

Boundary-layer control with unstructured uncertainties with application to adaptive autopilots

Control with unstructured uncertainties refers to controlling systems where not only the parameters are unknown, but also the way the parameters appear in the dynamics. This problem becomes pivotal in autopilots, where a unified control architecture is sought for aerial/ground/marine vehicles with different structures. By only making use of basic Euler-Lagrange properties valid in most mechanical systems independently of their specific structure, this brief proposes an adaptive design that does not rely on structural knowledge of the uncertainties. The proposed adaptive method, here validated in the ArduPlane module of ArduPilot, applies also to other modules like ArduCopter, ArduRover, ArduSub. Enhanced performance with respect to state-of-the-art methods addressing unstructured and state-dependent uncertainties is verified.

eess.SY

A Cooperative Implementation of Mesh Stability in Vehicular Platoons

This work studies the problem of mesh stability in connected and automated vehicles. Mesh stability, also known as 2D string stability, refers to studying how disturbances propagate in vehicular platoons in both longitudinal and lateral direction. As opposed to available decentralized results only relying on on-board sensing, the distinguishing feature of this work is a cooperative version of mesh stability, where onboard sensing is augmented by vehicle-to-vehicle communication. This cooperative version dramatically improves the state-of-the-art decentralized performance: for longitudinal control, a new cooperative non-identical protocol (i.e. with non-identical control gains) is proposed that improves the state-of-the-art decentralized non-identical protocol in terms of scalability of the control gains and strong notion of string stability. For lateral control, after showing that the non-identical approach is not necessary, a cooperative non-identical protocol is proposed to achieve another strong notion of string stability. Robustness of the proposed implementation against vehicle-to-vehicle communication delays and actuation time lags is considered. Numerical experiments, also performed with the vehicle simulator CarSim, validate the robustness and effectiveness of the proposed protocol.

eess.SY

Intriguing Electronic Structures of C8 and C12 Carbon Rings

We report on the ground and numerous excited electronic states. In the ground state the C4n rings are closed-shell systems possessing polyynic structures and can be classified as double anti-aromatic molecules. In their energetically lowest lying triplet state the rings exhibit aromatic cumulenic structures. The overall change in the electronic structures is rather dramatic upon the found moderate geometric changes from polyynic to cumulenic structure. Among others, Hund's rule is violated in both C8 and C12 in their cumulenic structures. We mention that until now, graphene is the only carbon allotrope reported to violate Hund's rule. The reasons for the violation are analyzed. Much effort has been invested to understand the relaxation pathways of the low-lying states leading the C8 from polyynic to cumulenic geometry and vice versa. On its minimum energy path, the first singlet excited state changes from open-shell character in the polyynic structure to a closed-shell state in the cumulenic structure. The cumulenic state lowest in energy is an open-shell singlet which relaxes to the closed-shell polyynic ground state.

physics.chem-ph

Towards Efficient Convolutional Neural Network for Embedded Hardware via Multi-Dimensional Pruning

In this paper, we propose TECO, a multi-dimensional pruning framework to collaboratively prune the three dimensions (depth, width, and resolution) of convolutional neural networks (CNNs) for better execution efficiency on embedded hardware. In TECO, we first introduce a two-stage importance evaluation framework, which efficiently and comprehensively evaluates each pruning unit according to both the local importance inside each dimension and the global importance across different dimensions. Based on the evaluation framework, we present a heuristic pruning algorithm to progressively prune the three dimensions of CNNs towards the optimal trade-off between accuracy and efficiency. Experiments on multiple benchmarks validate the advantages of TECO over existing state-of-the-art (SOTA) approaches. The code and pre-trained models are available at https://github.com/ntuliuteam/Teco.

cs.CV

On Hardware-Aware Design and Optimization of Edge Intelligence

Edge intelligence systems, the intersection of edge computing and artificial intelligence (AI), are pushing the frontier of AI applications. However, the complexity of deep learning models and heterogeneity of edge devices make the design of edge intelligence systems a challenging task. Hardware-agnostic methods face some limitations when implementing edge systems. Thus, hardware-aware methods are attracting more attention recently. In this paper, we present our recent endeavors in hardware-aware design and optimization for edge intelligence. We delve into techniques such as model compression and neural architecture search to achieve efficient and effective system designs. We also discuss some challenges in hardware-aware paradigm.

cs.AR

Collate: Collaborative Neural Network Learning for Latency-Critical Edge Systems

Federated Learning (FL) empowers multiple clients to collaboratively learn a model, enlarging the training data of each client for high accuracy while protecting data privacy. However, when deploying FL in real-time edge systems, the heterogeneity of devices among systems has a severe impact on the performance of the inferred model. Existing optimizations on FL focus on improving the training efficiency but fail to speed up inference, especially when there is a latency constraint. In this work, we propose Collate, a novel training framework that collaboratively learns heterogeneous models to meet the latency constraints of multiple edge systems simultaneously. We design a dynamic zeroizing-recovering method to adjust each local model architecture for high accuracy under its latency constraint. A proto-corrected federated aggregation scheme is also introduced to aggregate all heterogeneous local models, satisfying the latency constraint of different systems with only one training process and maintaining high accuracy. Extensive experiments indicate that, compared to state-of-the-art methods and under a latency constraint, our extended models can improve the accuracy by 1.96% on average, and our shrunk models can also obtain a 3.09% accuracy improvement on average, with almost no extra training overhead. The related codes and data will be available at https://github.com/ntuliuteam/Collate

cs.LG

FedTR: Federated Learning Framework with Transfer Learning for Industrial Visual Inspection

Federated learning (FL) is a collaborative learning scheme to train deep learning models, where collaborating parties can consolidate their models without sharing local data with other parties, hence preserving data privacy. Nevertheless, when implementing FL in Industrial visual inspection (IVI), the constraints posed by limited data availability and the intricate nature of the inspection tasks significantly impact the performance of the resulting model. This paper introduces FedTR, a novel FL framework incorporating transfer learning designed for Autonomous IVI, focusing on the challenging task of identifying label defects through end-to-end text recognition. Transfer learning is a method that leverages the knowledge of a pre-trained model to adapt to a different dataset. FedTR initially trains the model using a publicly available dataset, after which performs the essential federated learning process with model fine-tuning on the distributed and limited private data. Extensive experiment results demonstrate the effectiveness and feasibility of FedTR on private ink cartridge datasets for label defect identification. FedTR achieves an end-to-end text recognition word-level accuracy of 95.5% and 94.2% on homogeneous and heterogeneous data respectively. Additionally, it attains performance levels that are on par with those achieved through centralized training.

cs.CV

Latency-Constrained DNN Architecture Learning for Edge Systems using Zerorized Batch Normalization

Deep learning applications have been widely adopted on edge devices, to mitigate the privacy and latency issues of accessing cloud servers. Deciding the number of neurons during the design of a deep neural network to maximize performance is not intuitive. Particularly, many application scenarios are real-time and have a strict latency constraint, while conventional neural network optimization methods do not directly change the temporal cost of model inference for latency-critical edge systems. In this work, we propose a latency-oriented neural network learning method to optimize models for high accuracy while fulfilling the latency constraint. For efficiency, we also introduce a universal hardware-customized latency predictor to optimize this procedure to learn a model that satisfies the latency constraint by only a one-shot training process. The experiment results reveal that, compared to state-of-the-art methods, our approach can well-fit the 'hard' latency constraint and achieve high accuracy. Under the same training settings as the original model and satisfying a 34 ms latency constraint on the ImageNet-100 dataset, we reduce GoogLeNet's latency from 40.32 ms to 34 ms with a 0.14% accuracy reduction on the NVIDIA Jetson Nano. When coupled with quantization, our method can be further improved to only 0.04% drop for GoogLeNet. On the NVIDIA Jetson TX2, we compress VGG-19 from 119.98 ms to 34 ms and even improve its accuracy by 0.5%, and we scale GoogLeNet up from 20.27 ms to 34 ms and achieve higher accuracy by 0.78%. We also open source this framework at https://github.com/ntuliuteam/ZeroBN

cs.LG

EdgeCompress: Coupling Multidimensional Model Compression and Dynamic Inference for EdgeAI

Convolutional neural networks (CNNs) have demonstrated encouraging results in image classification tasks. However, the prohibitive computational cost of CNNs hinders the deployment of CNNs onto resource-constrained embedded devices. To address this issue, we propose EdgeCompress, a comprehensive compression framework to reduce the computational overhead of CNNs. In EdgeCompress, we first introduce dynamic image cropping (DIC), where we design a lightweight foreground predictor to accurately crop the most informative foreground object of input images for inference, which avoids redundant computation on background regions. Subsequently, we present compound shrinking (CS) to collaboratively compress the three dimensions (depth, width, and resolution) of CNNs according to their contribution to accuracy and model computation. DIC and CS together constitute a multidimensional CNN compression framework, which is able to comprehensively reduce the computational redundancy in both input images and neural network architectures, thereby improving the inference efficiency of CNNs. Further, we present a dynamic inference framework to efficiently process input images with different recognition difficulties, where we cascade multiple models with different complexities from our compression framework and dynamically adopt different models for different input images, which further compresses the computational redundancy and improves the inference efficiency of CNNs, facilitating the deployment of advanced CNNs onto embedded hardware. Experiments on ImageNet-1K demonstrate that EdgeCompress reduces the computation of ResNet-50 by 48.8% while improving the top-1 accuracy by 0.8%. Meanwhile, we improve the accuracy by 4.1% with similar computation compared to HRank, the state-of-the-art compression framework. The source code and models are available at https://github.com/ntuliuteam/edge-compress

cs.CV

Smart Scissor: Coupling Spatial Redundancy Reduction and CNN Compression for Embedded Hardware

Scaling down the resolution of input images can greatly reduce the computational overhead of convolutional neural networks (CNNs), which is promising for edge AI. However, as an image usually contains much spatial redundancy, e.g., background pixels, directly shrinking the whole image will lose important features of the foreground object and lead to severe accuracy degradation. In this paper, we propose a dynamic image cropping framework to reduce the spatial redundancy by accurately cropping the foreground object from images. To achieve the instance-aware fine cropping, we introduce a lightweight foreground predictor to efficiently localize and crop the foreground of an image. The finely cropped images can be correctly recognized even at a small resolution. Meanwhile, computational redundancy also exists in CNN architectures. To pursue higher execution efficiency on resource-constrained embedded devices, we also propose a compound shrinking strategy to coordinately compress the three dimensions (depth, width, resolution) of CNNs. Eventually, we seamlessly combine the proposed dynamic image cropping and compound shrinking into a unified compression framework, Smart Scissor, which is expected to significantly reduce the computational overhead of CNNs while still maintaining high accuracy. Experiments on ImageNet-1K demonstrate that our method reduces the computational cost of ResNet50 by 41.5% while improving the top-1 accuracy by 0.3%. Moreover, compared to HRank, the state-of-the-art CNN compression framework, our method achieves 4.1% higher top-1 accuracy at the same computational cost. The codes and data are available at https://github.com/ntuliuteam/smart-scissor

cs.CV

SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control

Geometry-conditioned 3D scene generation enables the creation of 3D environments from user-provided geometry, offering direct control over scene structure and object layout. To generate such 3D scenes, current methods commonly adopt a three-stage design that first defines a view schedule, then synthesizes multi-view observations along the scheduled views, and finally reconstructs a 3D representation from the generated images. However, defining the view schedule becomes a major bottleneck for outdoor scenes, where large, unstructured, and unbounded geometry makes it difficult to obtain views that provide sufficient coverage while supporting stable generation. To address this bottleneck, we present SceneFrom3D, a framework that automatically schedules views from outdoor input geometries. SceneFrom3D constructs a directed generation graph whose nodes represent anchor views and whose edges represent interpolation trajectories, defining which views to synthesize, which view pairs to interpolate, and in which order generation should proceed. Beyond automatic view scheduling, SceneFrom3D further improves controllability through object-level conditioning, assigning each object an identity image for appearance guidance and a geometry-adherence parameter for region-wise control over the input geometry. Experiments demonstrate that SceneFrom3D achieves state-of-the-art geometry-conditioned outdoor 3D scene generation, producing high-quality scenes with controllable object appearance and geometry adherence.

cs.GR

Multigrid Training for Molecular Generation using Graph Neural Networks

Deep learning has demonstrated significant success for modeling biochemical molecular systems, where inputs are commonly represented as graphs or 3D grids. A major challenge is that computational cost scales with resolution, making full graph/grid computation of molecular densities expensive and often unstable. We introduce a multigrid training strategy that leverages low-resolution optimization to accelerate learning at higher resolution through parameter transfer across discretizations. For graph molecular representations, we progressively transfer parameters learned from a coarse graph to a sequence of increasingly finer graphs via biased random walk upsampling. For 3D molecular generation, we voxelize the molecular structures at multiple resolutions, pretrain a coarse-resolution conditional Variational Autoencoder (CVAE), and initialize a fine-resolution CVAE by transferring shape compatible convolutional parameters from the coarse model. Numerical experiments on receptor-conditioned 3D Ligand generation show that multigrid training accelerates convergence and improves generalization compared to training from scratch.

cs.LG

AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference

As large language models scale to longer contexts, loading the growing KV cache during attention computation becomes a critical bottleneck. Previous work has shown that attention computation is dominated by a small subset of tokens. This motivates block sparse attention methods that partition the KV cache into fixed-size blocks and selectively compute attention over those blocks exhibiting high importance. However, these methods assign a uniform block size across all attention heads, implicitly assuming homogeneous behavior throughout the model. Our analysis reveals that this assumption is flawed: attention heads exhibit widely varying sensitivity to block granularity, and uniformity leads to suboptimal accuracy. We present AB-Sparse, a training-free algorithm-system co-designed framework that improves accuracy while preserving throughput. AB-Sparse introduces lightweight adaptive block size allocation across attention heads to improve accuracy. To compensate for the additional memory overhead, it further employs lossless block centroid quantization. In addition, custom GPU kernels are developed to support efficient execution with variable block sizes. Evaluation results demonstrate that AB-Sparse achieves an accuracy improvement of up to 5.43% over existing block sparse attention baselines without throughput overhead.

cs.DC

C2RustXW: Program-Structure-Aware C-to-Rust Translation via Program Analysis and LLM

The growing adoption of Rust for its memory safety and performance has increased the demand for effective migration of legacy C codebases. However, existing rule-based translators (e.g., \ctorust) often generate verbose, non-idiomatic code that preserves unsafe C semantics, limiting readability, maintainability, and practical adoption. Moreover, manual post-processing of such outputs is labor-intensive and rarely yields high-quality Rust code, posing a significant barrier to large-scale migration. To address these limitations, we present \tool, a program-structure-aware C-to-Rust translation approach that integrates program analysis with Large Language Models (LLMs). \tool extracts the multi-level program structure, including global symbols, function dependencies, and control- and data-flow information, and encodes these as structured textual representations injected into LLM prompts to guide translation and repair. Based on this design, \tool performs dependency-aware translation and adopts a multi-stage repair pipeline that combines rule-based and structure-guided LLM-based techniques to ensure syntactic correctness. For semantic correctness, \tool further integrates execution-based validation with structure-guided reasoning to localize and repair behavioral inconsistencies. Experimental results show that \tool achieves 100\% syntactic correctness on CodeNet and 97.78\% on GitHub, while significantly reducing code size (up to 43.70\%) and unsafe usage (to 5.75\%). At the project level, \tool achieves perfect syntactic correctness and an average semantic correctness of 78.87\%, demonstrating its effectiveness for practical and scalable C-to-Rust migration.

cs.SE