Skepxels: Spatio-temporal Image Representation of Human Skeleton Joints for Action Recognition

by   Jian Liu, et al.

Human skeleton joints are popular for action analysis since they can be easily extracted from videos to discard background noises. However, current skeleton representations do not fully benefit from machine learning with CNNs. We propose "Skepxels" a spatio-temporal representation for skeleton sequences to fully exploit the "local" correlations between joints using the 2D convolution kernels of CNN. We transform skeleton videos into images of flexible dimensions using Skepxels and develop a CNN-based framework for effective human action recognition using the resulting images. Skepxels encode rich spatio-temporal information about the skeleton joints in the frames by maximizing a unique distance metric, defined collaboratively over the distinct joint arrangements used in the skeletal image. Moreover, they are flexible in encoding compound semantic notions such as location and speed of the joints. The proposed action recognition exploits the representation in a hierarchical manner by first capturing the micro-temporal relations between the skeleton joints with the Skepxels and then exploiting their macro-temporal relations by computing the Fourier Temporal Pyramids over the CNN features of the skeletal images. We extend the Inception-ResNet CNN architecture with the proposed method and improve the state-of-the-art accuracy by 4.4 human activity dataset. On the medium-sized N-UCLA and UTH-MHAD datasets, our method outperforms the existing results by 5.7


page 4

page 7


Investigation of Different Skeleton Features for CNN-based 3D Action Recognition

Deep learning techniques are being used in skeleton based action recogni...

Skeleton based Activity Recognition by Fusing Part-wise Spatio-temporal and Attention Driven Residues

There exist a wide range of intra class variations of the same actions a...

Action Recognition with Visual Attention on Skeleton Images

Action recognition with 3D skeleton sequences is becoming popular due to...

A Fine-to-Coarse Convolutional Neural Network for 3D Human Action Recognition

This paper presents a new framework for human action recognition from 3D...

Spatio-Temporal LSTM with Trust Gates for 3D Human Action Recognition

3D action recognition - analysis of human actions based on 3D skeleton d...

Real-Time Driver State Monitoring Using a CNN Based Spatio-Temporal Approach

Many road accidents occur due to distracted drivers. Today, driver monit...

Toward Accurate Person-level Action Recognition in Videos of Crowded Scenes

Detecting and recognizing human action in videos with crowded scenes is ...

Please sign up or login with your details

Forgot password? Click here to reset