arXiv ScienceSearch

arXiv · 2407.10975

Stream State-tying for Sign Language Recognition

Abstract

In this paper, a novel approach to sign language recognition based on state tying in each of data streams is presented. In this framework, it is assumed that hand gesture signal is represented in terms of six synchronous data streams, i.e., the left/right hand position, left/right hand orientation and left/right handshape. This approach offers a very accurate representation of the sign space and keeps the number of parameters reasonably small in favor of a fast decoding. Experiments were carried out for 5177 Chinese signs. The real time isolated recognition rate is 94.8%. For continuous sign recognition, the word correct rate is 91.4%. Keywords: Sign language recognition; Automatic sign language translation; Hand gesture recognition; Hidden Markov models; State-tying; Multimodal user interface; Virtual reality; Man-machine systems.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jiyong Ma, Wen Gao, Chunli Wang. 2024-04-21. Stream State-tying for Sign Language Recognition. https://arxiv.org/abs/2407.10975

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Bidirectional Temporal Dynamics Modeling for EEG-based Driving Fatigue Recognition

Driving fatigue is a major contributor to traffic accidents and poses a serious threat to road safety. Electroencephalography (EEG) provides a direct measurement of neural activity, yet EEG-based fatigue recognition is hindered by strong non-stationarity and asymmetric neural dynamics. To address these challenges, we propose DeltaGateNet, a novel framework that explicitly captures Bidirectional temporal dynamics for EEG-based driving fatigue recognition. Our key idea is to introduce a Bidirectional Delta module that decomposes first-order temporal differences into positive and negative components, enabling explicit modeling of asymmetric neural activation and suppression patterns. Furthermore, we design a Gated Temporal Convolution module to capture long-term temporal dependencies for each EEG channel using depthwise temporal convolutions and residual learning, preserving channel-wise specificity while enhancing temporal representation robustness. Extensive experiments conducted under both intra-subject and inter-subject evaluation settings on the public SEED-VIG and SADT driving fatigue datasets demonstrate that DeltaGateNet consistently outperforms existing methods. On SEED-VIG, DeltaGateNet achieves an intra-subject accuracy of 81.89% and an inter-subject accuracy of 55.55%. On the balanced SADT 2022 dataset, it attains intra-subject and inter-subject accuracies of 96.81% and 83.21%, respectively, while on the unbalanced SADT 2952 dataset, it achieves 96.84% intra-subject and 84.49% inter-subject accuracy. These results indicate that explicitly modeling Bidirectional temporal dynamics yields robust and generalizable performance under varying subject and class-distribution conditions.

cs.OH

Fixing ill-formed UTF-16 strings with SIMD instructions

UTF-16 is a widely used Unicode encoding representing characters with one or two 16-bit code units. The format relies on surrogate pairs to encode characters beyond the Basic Multilingual Plane, requiring a high surrogate followed by a low surrogate. Ill-formed UTF-16 strings -- where surrogates are mismatched -- can arise from data corruption or improper encoding, posing security and reliability risks. Consequently, programming languages such as JavaScript include functions to fix ill-formed UTF-16 strings by replacing mismatched surrogates with the Unicode replacement character (U+FFFD). We propose using Single Instruction, Multiple Data (SIMD) instructions to handle multiple code units in parallel, enabling faster and more efficient execution. Our software is part of the Google JavaScript engine (V8) and thus part of several major Web browsers.

cs.OH

MRSeqStudio: MRI Sequence Design and Simulation as a Service in a Free and Open-Source Web Platform

MRI sequence prototyping increasingly relies on graphical design environments and numerical simulators to accelerate development and validation. While several platforms support interactive sequence construction, fully web-based solutions that combine integrated phantom management, high-fidelity Bloch simulation, and scalable multi-user deployment remain limited. We present MRSeqStudio, a web-based platform for interactive MR sequence design and simulation. The tool adopts a block-based representation model with real-time visualization and native JSON/Pulseq export. Simulations are performed using the GPU-enabled Bloch simulator KomaMRI, which enables accurate modeling of arbitrary pulse sequences and phantoms within an installation-free architecture. The system separates front-end interaction from back-end simulation services to support concurrent multi-user access. Sequence validity was assessed by comparing GRE and bSSFP implementations against equivalent sequences designed in mtrk and gammaSTAR. The resulting images showed minimal absolute differences and high mean structural similarity indices (SSIM). Stress testing under burst-request conditions demonstrated stable performance with up to 100 concurrent users on a high-performance desktop deployment. A comparative workflow analysis with mtrk and gammaSTAR further examined differences in representation models, parameter propagation strategies, and integration levels across platforms, highlighting the relative strengths and limitations of each tool. Results indicate that MRSeqStudio provides a reliable and accessible environment for MR sequence prototyping, combining web-native deployment with Bloch-level simulation fidelity and integrated phantom visualization.

cs.OH