Learning 3D Semantic Segmentation with only 2D Image Supervision

by   Kyle Genova, et al.

With the recent growth of urban mapping and autonomous driving efforts, there has been an explosion of raw 3D data collected from terrestrial platforms with lidar scanners and color cameras. However, due to high labeling costs, ground-truth 3D semantic segmentation annotations are limited in both quantity and geographic diversity, while also being difficult to transfer across sensors. In contrast, large image collections with ground-truth semantic segmentations are readily available for diverse sets of scenes. In this paper, we investigate how to use only those labeled 2D image collections to supervise training 3D semantic segmentation models. Our approach is to train a 3D model from pseudo-labels derived from 2D semantic image segmentations using multiview fusion. We address several novel issues with this approach, including how to select trusted pseudo-labels, how to sample 3D scenes with rare object categories, and how to decouple input features from 2D images from pseudo-labels during training. The proposed network architecture, 2D3DNet, achieves significantly better performance (+6.2-11.4 mIoU) than baselines during experiments on a new urban dataset with lidar and images captured in 20 cities across 5 continents.


page 1

page 3

page 4

page 7

page 11

page 12


Semantic Segmentation on Swiss3DCities: A Benchmark Study on Aerial Photogrammetric 3D Pointcloud Dataset

We introduce a new outdoor urban 3D pointcloud dataset, covering a total...

PillarSegNet: Pillar-based Semantic Grid Map Estimation using Sparse LiDAR Data

Semantic understanding of the surrounding environment is essential for a...

Drive Segment: Unsupervised Semantic Segmentation of Urban Scenes via Cross-modal Distillation

This work investigates learning pixel-wise semantic image segmentation i...

Less is More: Reducing Task and Model Complexity for 3D Point Cloud Semantic Segmentation

Whilst the availability of 3D LiDAR point cloud data has significantly g...

A Bayesian Approach to Reinforcement Learning of Vision-Based Vehicular Control

In this paper, we present a state-of-the-art reinforcement learning meth...

Can Ground Truth Label Propagation from Video help Semantic Segmentation?

For state-of-the-art semantic segmentation task, training convolutional ...

Urban Scene Semantic Segmentation with Low-Cost Coarse Annotation

For best performance, today's semantic segmentation methods use large an...

Please sign up or login with your details

Forgot password? Click here to reset