The 7th International Conference on Next Generation Computing 2021 (2021.11)바로가기
페이지
pp.50-52
저자
Zongjing Cao, Yan Li, Byeong-Seok Shin
언어
영어(ENG)
URL
https://www.earticle.net/Article/A448006
원문정보
초록
영어
Recently, hand gesture recognition based on deep 3D convolution neural networks has made great progress. However, the large number of weight parameters that need to be optimized leads to its expensive computational cost. We introduced a transformer-based framework for hand gesture recognition, which is a fully self-attentional architecture. The framework abandons the conventional methods that rely on 3D convolution and proposed an approach to classify actions by focusing on the entire video sequence. In addition, we use a lightweight hand detector to continuously sample the video only when a hand is detected in the video sequence, thus reducing the computational consumption of the system. Experiments on two human hand gesture recognition benchmark datasets show the superiority of the proposed method, compared with existing state-of-the-art methods.
목차
Abstract I. INTRODUCTION II. METHODOLOGY III. EXPERIMENTS ACKNOWLEDGMENT REFERENCES