arXiv ScienceSearch

arXiv · 2404.15279

Jointly Modeling Spatio-Temporal Features of Tactile Signals for Action Classification

Abstract

Tactile signals collected by wearable electronics are essential in modeling and understanding human behavior. One of the main applications of tactile signals is action classification, especially in healthcare and robotics. However, existing tactile classification methods fail to capture the spatial and temporal features of tactile signals simultaneously, which results in sub-optimal performances. In this paper, we design Spatio-Temporal Aware tactility Transformer (STAT) to utilize continuous tactile signals for action classification. We propose spatial and temporal embeddings along with a new temporal pretraining task in our model, which aims to enhance the transformer in modeling the spatio-temporal features of tactile signals. Specially, the designed temporal pretraining task is to differentiate the time order of tubelet inputs to model the temporal properties explicitly. Experimental results on a public action classification dataset demonstrate that our model outperforms state-of-the-art methods in all metrics.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jimmy Lin, Junkai Li, Jiasi Gao, Weizhi Ma, Yang Liu. 2024-01-21. Jointly Modeling Spatio-Temporal Features of Tactile Signals for Action Classification. https://arxiv.org/abs/2404.15279

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Robust Eavesdropping in the Presence of Adversarial Communications for RF Fingerprinting

Deep learning is an effective approach for performing radio frequency (RF) fingerprinting, which aims to identify the transmitter corresponding to received RF signals. However, beyond the intended receiver, malicious eavesdroppers can also intercept signals and attempt to fingerprint transmitters communicating over a wireless channel. Recent studies suggest that transmitters can counter such threats by embedding deep learning-based transferable adversarial attacks in their signals before transmission. In this work, we develop a time-frequency-based eavesdropper architecture that is capable of withstanding such transferable adversarial perturbations and thus able to perform effective RF fingerprinting. We theoretically demonstrate that adversarial perturbations injected by a transmitter are confined to specific time-frequency regions that are insignificant during inference, directly increasing fingerprinting accuracy on perturbed signals intercepted by the eavesdropper. Empirical evaluations on a real-world dataset validate our theoretical findings, showing that deep learning-based RF fingerprinting eavesdroppers can achieve classification performance comparable to the intended receiver, despite efforts made by the transmitter to deceive the eavesdropper. Our framework reveals that relying on transferable adversarial attacks may not be sufficient to prevent eavesdroppers from successfully fingerprinting transmissions in next-generation deep learning-based communications systems.

eess.SP

Channel Estimation for Movable Antenna Systems: Challenges, Solutions, and Opportunities

Movable antenna (MA) has emerged as a promising technology for future wireless networks by exploiting channel variation over local antenna movement regions. However, accurate and efficient channel acquisition in MA systems remains challenging due to the trade-off between estimation accuracy and computational complexity. In this article, the MA channel model and the associated estimation framework are first reviewed, where channel information over the movement region is reconstructed from finite measurements by exploiting shared path parameters. The structured dependence of MA observations across space, time, and frequency naturally motivates the adoption of tensor-based modeling for channel estimation. Subsequently, we discuss the tensor-based signal model and corresponding parameter estimation methods from multidimensional observations. These methods are further compared with conventional channel estimation methods in terms of estimation accuracy, computational complexity, and general applicability. Furthermore, a representative case study is provided to illustrate the performance and characteristics of different algorithms under MA channel estimation settings. Finally, some future research directions for tensor decomposition-based channel estimation in MA systems are outlined.

eess.SP

Dense and Query-Set Prediction for J-Peak Detection in Pillow-Based Ballistocardiography

Pillow-based ballistocardiography (BCG) enables unobtrusive cardiac monitoring, but J-peak detection is commonly learned as dense sequence labeling although the desired output is a sparse event set. This letter asks a narrower question: how do dense and query-set outputs differ when the data, encoder, validation protocol, and event evaluator are controlled? We formulate one-dimensional point-set prediction with 64 learned queries and Hungarian assignment, and compare it with a shared-backbone dense Transformer and a U-Net--BiLSTM. Evaluation uses five-subject leave-one-subject-out testing, three independently seeded runs, validation-only postprocessing selection, and strict one-to-one peak association. The dense Transformer attains the highest subject-wise pooled F1 ($0.786\pm0.105$) and precision ($0.817\pm0.089$), whereas U-Net--BiLSTM obtains $0.780\pm0.094$ F1. Set+DN reaches $0.764\pm0.112$ F1 but the lowest beat-count error ($0.862\pm0.584$ beats/epoch), compared with $2.790\pm1.356$ for the dense Transformer. DN changes set-model F1 by only $+0.004$. The results identify distinct event-accuracy and count-fidelity operating points; they do not establish universal superiority of either output formulation.

eess.SP