XFormer: Fast and Accurate Monocular 3D Body Capture

05/18/2023
by   Lihui Qian, et al.
0

We present XFormer, a novel human mesh and motion capture method that achieves real-time performance on consumer CPUs given only monocular images as input. The proposed network architecture contains two branches: a keypoint branch that estimates 3D human mesh vertices given 2D keypoints, and an image branch that makes predictions directly from the RGB image features. At the core of our method is a cross-modal transformer block that allows information to flow across these two branches by modeling the attention between 2D keypoint coordinates and image spatial features. Our architecture is smartly designed, which enables us to train on various types of datasets including images with 2D/3D annotations, images with 3D pseudo labels, and motion capture datasets that do not have associated images. This effectively improves the accuracy and generalization ability of our system. Built on a lightweight backbone (MobileNetV3), our method runs blazing fast (over 30fps on a single CPU core) and still yields competitive accuracy. Furthermore, with an HRNet backbone, XFormer delivers state-of-the-art performance on Huamn3.6 and 3DPW datasets.

READ FULL TEXT

page 5

page 11

page 13

page 14

research
07/25/2019

Cross Attention Network for Semantic Segmentation

In this paper, we address the semantic segmentation task with a deep net...
research
04/27/2021

KAMA: 3D Keypoint Aware Body Mesh Articulation

We present KAMA, a 3D Keypoint Aware Mesh Articulation approach that all...
research
09/01/2023

Fusing Monocular Images and Sparse IMU Signals for Real-time Human Motion Capture

Either RGB images or inertial signals have been used for the task of mot...
research
12/11/2020

Monocular Real-time Full Body Capture with Inter-part Correlations

We present the first method for real-time full body capture that estimat...
research
12/18/2020

Human 3D keypoints via spatial uncertainty modeling

We introduce a technique for 3D human keypoint estimation that directly ...
research
03/11/2023

DECOMPL: Decompositional Learning with Attention Pooling for Group Activity Recognition from a Single Volleyball Image

Group Activity Recognition (GAR) aims to detect the activity performed b...
research
09/13/2022

CAIBC: Capturing All-round Information Beyond Color for Text-based Person Retrieval

Given a natural language description, text-based person retrieval aims t...

Please sign up or login with your details

Forgot password? Click here to reset