arXiv ScienceSearch

arXiv subjects

Kan Wang

Publications and source records attributed to Kan Wang.

At least 19 recordsLinked to original sources

An Adaptive Environment-Aware Transformer Autoencoder for UAV-FSO with Dynamic Complexity Control

The rise of sixth-generation (6G) wireless networks sets high demands on UAV-assisted Free Space Optical (FSO) communications, where the channel environment becomes more complex and variable due to both atmospheric turbulence and UAV-induced vibrations. These factors increase the challenge of maintaining reliable communication and require adaptive processing methods. Autoencoders are promising as they learn optimal encodings from channel data. However, existing autoencoder designs are generic and lack the specific adaptability and computational flexibility needed for UAV-FSO scenarios. To address this, we propose AEAT-AE (Adaptive Environment-aware Transformer Autoencoder), a Transformer-based framework that integrates environmental parameters into both encoder and decoder via a cross-attention mechanism. Moreover, AEAT-AE incorporates a Deep Q-Network (DQN) that dynamically selects which layers of the Transformer autoencoder to activate based on real-time environmental inputs, effectively balancing performance and computational cost. Simulation results demonstrate that AEAT-AE outperforms conventional methods in bit error rate while maintaining efficient runtime, representing a novel tailored solution for next-generation UAV-FSO communications.

eess.SY

HS-SLAM: A Fast and Hybrid Strategy-Based SLAM Approach for Low-Speed Autonomous Driving

Visual-inertial simultaneous localization and mapping (SLAM) is a key module of robotics and low-speed autonomous vehicles, which is usually limited by the high computation burden for practical applications. To this end, an innovative strategy-based hybrid framework HS-SLAM is proposed to integrate the advantages of direct and feature-based methods for fast computation without decreasing the performance. It first estimates the relative positions of consecutive frames using IMU pose estimation within the tracking thread. Then, it refines these estimates through a multi-layer direct method, which progressively corrects the relative pose from coarse to fine, ultimately achieving accurate corner-based feature matching. This approach serves as an alternative to the conventional constant-velocity tracking model. By selectively bypassing descriptor extraction for non-critical frames, HS-SLAM significantly improves the tracking speed. Experimental evaluations on the EuRoC MAV dataset demonstrate that HS-SLAM achieves higher localization accuracies than ORB-SLAM3 while improving the average tracking efficiency by 15%.

cs.RO

This Time is Different: An Observability Perspective on Time Series Foundation Models

We introduce Toto, a time series forecasting foundation model with 151 million parameters. Toto uses a modern decoder-only architecture coupled with architectural innovations designed to account for specific challenges found in multivariate observability time series data. Toto's pre-training corpus is a mixture of observability data, open datasets, and synthetic data, and is 4-10$\times$ larger than those of leading time series foundation models. Additionally, we introduce BOOM, a large-scale benchmark consisting of 350 million observations across 2,807 real-world time series. For both Toto and BOOM, we source observability data exclusively from Datadog's own telemetry and internal observability metrics. Extensive evaluations demonstrate that Toto achieves state-of-the-art performance on both BOOM and on established general purpose time series forecasting benchmarks. Toto's model weights, inference code, and evaluation scripts, as well as BOOM's data and evaluation code, are all available as open source under the Apache 2.0 License available at https://huggingface.co/Datadog/Toto-Open-Base-1.0 and https://github.com/DataDog/toto.

cs.LG

Artificial Intelligence in Reactor Physics: Current Status and Future Prospects

Reactor physics is the study of neutron properties, focusing on using models to examine the interactions between neutrons and materials in nuclear reactors. Artificial intelligence (AI) has made significant contributions to reactor physics, e.g., in operational simulations, safety design, real-time monitoring, core management and maintenance. This paper presents a comprehensive review of AI approaches in reactor physics, especially considering the category of Machine Learning (ML), with the aim of describing the application scenarios, frontier topics, unsolved challenges and future research directions. From equation solving and state parameter prediction to nuclear industry applications, this paper provides a step-by-step overview of ML methods applied to steady-state, transient and combustion problems. Most literature works achieve industry-demanded models by enhancing the efficiency of deterministic methods or correcting uncertainty methods, which leads to successful applications. However, research on ML methods in reactor physics is somewhat fragmented, and the ability to generalize models needs to be strengthened. Progress is still possible, especially in addressing theoretical challenges and enhancing industrial applications such as building surrogate models and digital twins.

cs.LG

Structure of an axisymmetric turbulent boundary layer under adverse pressure gradient: a large-eddy simulation study

The spatial characteristics and structure of an axisymmetric turbulent boundary layer under strong adverse pressure gradient and weak transverse curvature are investigated using incompressible large-eddy simulation. The boundary layer is on a $20^{\circ}$ tail cone of a body of revolution at a length-based Reynolds number of $1.9\times10^6$. The simulation results are in agreement with the experimental measurements of Balantrapu et al. (J. Fluid Mech., vol. 929, 2021) and significantly expand the experimental results with new flow-field details and physical insights. The mean streamwise velocity profiles exhibit a shortened logarithmic region and a longer wake region compared with planar boundary layers at zero pressure gradient. With the embedded-shear-layer scaling, self-similarity is observed for the mean velocity and all three components of turbulence intensity. The azimuthal-wavenumber spectra of streamwise velocity fluctuations possess two peaks in the wall-normal direction, an inner peak at the wavelength of approximately 100 wall units and an outer peak in the wake region with growing strength, wavelength and distance to the wall in the downstream direction. Two-point correlations of streamwise velocity fluctuations show significant downstream growth and elongation of turbulence structures with increasing inclination angle. However, relative to the boundary-layer thickness, the correlation structures decrease in size in the downstream direction. The distributions of streamwise and wall-normal integral lengths across the boundary-layer thickness resemble those of zero-pressure-gradient planar boundary layers, whereas the azimuthal integral length deviates from the planar boundary-layer behavior at downstream stations.

physics.flu-dyn

Toto: Time Series Optimized Transformer for Observability

This technical report describes the Time Series Optimized Transformer for Observability (Toto), a new state of the art foundation model for time series forecasting developed by Datadog. In addition to advancing the state of the art on generalized time series benchmarks in domains such as electricity and weather, this model is the first general-purpose time series forecasting foundation model to be specifically tuned for observability metrics. Toto was trained on a dataset of one trillion time series data points, the largest among all currently published time series foundation models. Alongside publicly available time series datasets, 75% of the data used to train Toto consists of fully anonymous numerical metric data points from the Datadog platform. In our experiments, Toto outperforms existing time series foundation models on observability data. It does this while also excelling at general-purpose forecasting tasks, achieving state-of-the-art zero-shot performance on multiple open benchmark datasets.

cs.LG

Extension of all-optical reconstruction method for isolated attosecond pulses using high-harmonic generation streaking spectra

An all-optical method for directly reconstructing the spectral phase of isolated attosecond pulse (IAP) has been proposed recently [New J. Phys. 25, 083003 (2023)]. This method is based on the high-harmonic generation (HHG) streaking spectra generated by an IAP and a time-delayed intense infrared (IR) laser, which can be accurately simulated by an extended quantitative rescattering model. Here we extend the retrieval algorithm in this method to successfully retrieve the spectral phase of an shaped IAP, which has a spectral minimum, a phase jump about $\pi$, and a "split" temporal profile. We then reconstruct the carrier-envelope phase of IR laser from HHG streaking spectra. And we finally discuss the retrieval of the phase of high harmonics by the intense IR laser alone using the Fourier transform of HHG streaking spectra.

physics.optics

Enhancing Resource Utilization of Non-terrestrial Networks Using Temporal Graph-based Deterministic Routing

Deterministic routing has emerged as a promising technology for future non-terrestrial networks (NTNs), offering the potential to enhance service performance and optimize resource utilization. However, the dynamic nature of network topology and resources poses challenges in establishing deterministic routing. These challenges encompass the intricacy of jointly scheduling transmission links and cycles, as well as the difficulty of maintaining stable end-to-end (E2E) routing paths. To tackle these challenges, our work introduces an efficient temporal graph-based deterministic routing strategy. Initially, we utilize a time-expanded graph (TEG) to represent the heterogeneous resources of an NTN in a time-slotted manner. With TEG, we meticulously define each necessary constraint and formulate the deterministic routing problem. Subsequently, we transform this nonlinear problem equivalently into solvable integer linear programming (ILP), providing a robust yet time-consuming performance upper bound. To address the considered problem with reduced complexity, we extend TEG by introducing virtual nodes and edges. This extension facilitates a uniform representation of heterogeneous network resources and traffic transmission requirements. Consequently, we propose a polynomial-time complexity algorithm, enabling the dynamic selection of optimal transmission links and cycles on a hop-by-hop basis. Simulation results validate that the proposed algorithm yields significant performance gains in traffic acceptance, justifying its additional complexity compared to existing routing strategies.

cs.NI

Context Sensing Attention Network for Video-based Person Re-identification

Video-based person re-identification (ReID) is challenging due to the presence of various interferences in video frames. Recent approaches handle this problem using temporal aggregation strategies. In this work, we propose a novel Context Sensing Attention Network (CSA-Net), which improves both the frame feature extraction and temporal aggregation steps. First, we introduce the Context Sensing Channel Attention (CSCA) module, which emphasizes responses from informative channels for each frame. These informative channels are identified with reference not only to each individual frame, but also to the content of the entire sequence. Therefore, CSCA explores both the individuality of each frame and the global context of the sequence. Second, we propose the Contrastive Feature Aggregation (CFA) module, which predicts frame weights for temporal aggregation. Here, the weight for each frame is determined in a contrastive manner: i.e., not only by the quality of each individual frame, but also by the average quality of the other frames in a sequence. Therefore, it effectively promotes the contribution of relatively good frames. Extensive experimental results on four datasets show that CSA-Net consistently achieves state-of-the-art performance.

cs.CV

Metasurface Near-field Measurements with Incident Field Reconstruction using a Single Horn Antenna

A simple method of superimposing multiple near field scans using a single horn antenna in different configurations to characterize a planar electromagnetic metasurface is proposed and numerically demonstrated. It can be used to construct incident fields for which the metasurface is originally designed for, which may otherwise be difficult or not possible to achieve in practice. While this method involves additional effort by requiring multiple scans, it also provides flexibility for the incident field to be generated, simply by changing the objective of a numerical optimization which is used to find the required horn configurations for the different experiments. The proposed method is applicable to all linear time-invariant metasurfaces including space-time modulated structures.

physics.class-ph

Batch Coherence-Driven Network for Part-aware Person Re-Identification

Existing part-aware person re-identification methods typically employ two separate steps: namely, body part detection and part-level feature extraction. However, part detection introduces an additional computational cost and is inherently challenging for low-quality images. Accordingly, in this work, we propose a simple framework named Batch Coherence-Driven Network (BCD-Net) that bypasses body part detection during both the training and testing phases while still learning semantically aligned part features. Our key observation is that the statistics in a batch of images are stable, and therefore that batch-level constraints are robust. First, we introduce a batch coherence-guided channel attention (BCCA) module that highlights the relevant channels for each respective part from the output of a deep backbone model. We investigate channelpart correspondence using a batch of training images, then impose a novel batch-level supervision signal that helps BCCA to identify part-relevant channels. Second, the mean position of a body part is robust and consequently coherent between batches throughout the training process. Accordingly, we introduce a pair of regularization terms based on the semantic consistency between batches. The first term regularizes the high responses of BCD-Net for each part on one batch in order to constrain it within a predefined area, while the second encourages the aggregate of BCD-Nets responses for all parts covering the entire human body. The above constraints guide BCD-Net to learn diverse, complementary, and semantically aligned part-level features. Extensive experimental results demonstrate that BCDNet consistently achieves state-of-the-art performance on four large-scale ReID benchmarks.

cs.CV

A calibration-free method for biosensing in cell manufacturing

Chimeric antigen receptor T cell therapy has demonstrated innovative therapeutic effectiveness in fighting cancers; however, it is extremely expensive due to the intrinsic patient-to-patient variability in cell manufacturing. We propose in this work a novel calibration-free statistical framework to effectively recover critical quality attributes under the patient-to-patient variability. Specifically, we model this variability via a patient-specific calibration parameter, and use readings from multiple biosensors to construct a patient-invariance statistic, thereby alleviating the effect of the calibration parameter. A carefully formulated optimization problem and an algorithmic framework are presented to find the best patient-invariance statistic and the model parameters. Using the patient-invariance statistic, we can recover the critical quality attribute of interest, free from the calibration parameter. We demonstrate improvements of the proposed calibration-free method in different simulation experiments. In the cell manufacturing case study, our method not only effectively recovers viable cell concentration for monitoring, but also reveals insights for the cell manufacturing process.

q-bio.QM

Multi-task Learning with Coarse Priors for Robust Part-aware Person Re-identification

Part-level representations are important for robust person re-identification (ReID), but in practice feature quality suffers due to the body part misalignment problem. In this paper, we present a robust, compact, and easy-to-use method called the Multi-task Part-aware Network (MPN), which is designed to extract semantically aligned part-level features from pedestrian images. MPN solves the body part misalignment problem via multi-task learning (MTL) in the training stage. More specifically, it builds one main task (MT) and one auxiliary task (AT) for each body part on the top of the same backbone model. The ATs are equipped with a coarse prior of the body part locations for training images. ATs then transfer the concept of the body parts to the MTs via optimizing the MT parameters to identify part-relevant channels from the backbone model. Concept transfer is accomplished by means of two novel alignment strategies: namely, parameter space alignment via hard parameter sharing and feature space alignment in a class-wise manner. With the aid of the learned high-quality parameters, MTs can independently extract semantically aligned part-level features from relevant channels in the testing stage. MPN has three key advantages: 1) it does not need to conduct body part detection in the inference stage; 2) its model is very compact and efficient for both training and testing; 3) in the training stage, it requires only coarse priors of body part locations, which are easy to obtain. Systematic experiments on four large-scale ReID databases demonstrate that MPN consistently outperforms state-of-the-art approaches by significant margins. Code is available at https://github.com/WangKan0128/MPN.

cs.CV

CDPM: Convolutional Deformable Part Models for Semantically Aligned Person Re-identification

Part-level representations are essential for robust person re-identification. However, common errors that arise during pedestrian detection frequently result in severe misalignment problems for body parts, which degrade the quality of part representations. Accordingly, to deal with this problem, we propose a novel model named Convolutional Deformable Part Models (CDPM). CDPM works by decoupling the complex part alignment procedure into two easier steps: first, a vertical alignment step detects each body part in the vertical direction, with the help of a multi-task learning model; second, a horizontal refinement step based on attention suppresses the background information around each detected body part. Since these two steps are performed orthogonally and sequentially, the difficulty of part alignment is significantly reduced. In the testing stage, CDPM is able to accurately align flexible body parts without any need for outside information. Extensive experimental results demonstrate the effectiveness of the proposed CDPM for part alignment. Most impressively, CDPM achieves state-of-the-art performance on three large-scale datasets: Market-1501, DukeMTMC-ReID,and CUHK03.

cs.CV

Active Image Synthesis for Efficient Labeling

The great success achieved by deep neural networks attracts increasing attention from the manufacturing and healthcare communities. However, the limited availability of data and high costs of data collection are the major challenges for the applications in those fields. We propose in this work AISEL, an active image synthesis method for efficient labeling to improve the performance of the small-data learning tasks. Specifically, a complementary AISEL dataset is generated, with labels actively acquired via a physics-based method to incorporate underlining physical knowledge at hand. An important component of our AISEL method is the bidirectional generative invertible network (GIN), which can extract interpretable features from the training images and generate physically meaningful virtual images. Our AISEL method then efficiently samples virtual images not only further exploits the uncertain regions, but also explores the entire image space. We then discuss the interpretability of GIN both theoretically and experimentally, demonstrating clear visual improvements over the benchmarks. Finally, we demonstrate the effectiveness of our AISEL framework on aortic stenosis application, in which our method lower the labeling cost by $90\%$ while achieving a $15\%$ improvement in prediction accuracy.

cs.CV

Generative Invertible Networks (GIN): Pathophysiology-Interpretable Feature Mapping and Virtual Patient Generation

Machine learning methods play increasingly important roles in pre-procedural planning for complex surgeries and interventions. Very often, however, researchers find the historical records of emerging surgical techniques, such as the transcatheter aortic valve replacement (TAVR), are highly scarce in quantity. In this paper, we address this challenge by proposing novel generative invertible networks (GIN) to select features and generate high-quality virtual patients that may potentially serve as an additional data source for machine learning. Combining a convolutional neural network (CNN) and generative adversarial networks (GAN), GIN discovers the pathophysiologic meaning of the feature space. Moreover, a test of predicting the surgical outcome directly using the selected features results in a high accuracy of 81.55%, which suggests little pathophysiologic information has been lost while conducting the feature selection. This demonstrates GIN can generate virtual patients not only visually authentic but also pathophysiologically interpretable.

cs.CV

The Impact of Antenna Height Difference on the Performance of Downlink Cellular Networks

Capable of significantly reducing cell size and enhancing spatial reuse, network densification is shown to be one of the most dominant approaches to expand network capacity. Due to the scarcity of available spectrum resources, nevertheless, the over-deployment of network infrastructures, e.g., cellular base stations (BSs), would strengthen the inter-cell interference as well, thus in turn deteriorating the system performance. On this account, we investigate the performance of downlink cellular networks in terms of user coverage probability (CP) and network spatial throughput (ST), aiming to shed light on the limitation of network densification. Notably, it is shown that both CP and ST would be degraded and even diminish to be zero when BS density is sufficiently large, provided that practical antenna height difference (AHD) between BSs and users is involved to characterize pathloss. Moreover, the results also reveal that the increase of network ST is at the expense of the degradation of CP. Therefore, to balance the tradeoff between user and network performance, we further study the critical density, under which ST could be maximized under the CP constraint. Through a special case study, it follows that the critical density is inversely proportional to the square of AHD. The results in this work could provide helpful guideline towards the application of network densification in the next-generation wireless networks.

cs.IT

Information-Centric Wireless Networks with Virtualization and D2D Communications

Wireless network virtualization and information-centric networking (ICN) are two promising technologies for next generation wireless networks. Although some excellent works have focused on these two technologies, device-to-device (D2D) communications have not beeen investigated in information-centric virtualized cellular networks. Meanwhile, content caching in mobile devices has attracted great attentions due to the saved backhaul consumption or reduced transmission latency in D2D-assisted cellular networks. However, when it comes to the multi-operator scenario, the direct content sharing between different operators via D2D communications is typically infeasible. In this article, we propose a novel information-centric virtualized cellular network framework with D2D communications, enabling not only content caching in the air, but also inter-operator content sharing between mobile devices. Moreover, we describe the key components in the proposed framework, and present the interactions among them. In addition, we incorporate and formulate the content caching strategies in resource allocation optimization, to maximize the total utility of mobile virtual network operators (MVNOs) through caching popular contents in mobile devices. Simulations results demonstrate the effectiveness of the proposed framework and scheme with different system parameters.

cs.NI