arXiv · 2609.32244
Two-Stage Multi-View Gait Recognition with a Re-Embedding Network
Abstract
Gait recognition always remains challenging due to severe overfitting and the rigid view constraints common in single-stage approaches. We propose a two-stage framework, termed Translate-First-Then-Reason (TFTR), to address these issues. In the first stage, a shallow Siamese convolutional network with triplet loss maps Gait Energy Images (GEIs) into a 128-dimensional view-specific embedding space. In the second stage, these per-view embeddings are treated as tokens and processed by a 12-layer Transformer encoder, which re-projects them into a new space with improved cosine separability. This design enables flexible fusion of an arbitrary number of views at inference, overcoming the fixed-input limitations of prior methods. Trained on the OU-MVLP dataset (6,000 subjects) and evaluated on unseen CASIA-B across normal, bag-carrying, and coat-wearing conditions, our pipeline achieves 96.91\% single-view and 99.49\% three-view accuracy on OU-MVLP, and attains 100\% accuracy on CASIA-B with three views.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Long Hoang Le, Trung Thanh Ngo. 2026-09-26. Two-Stage Multi-View Gait Recognition with a Re-Embedding Network. https://arxiv.org/abs/2609.32244
Cite the original work for its findings. Save a collection to share your selection of sources.