Siamese Network for RGB-D Salient Object Detection and Beyond

by   Keren Fu, et al.

Existing RGB-D salient object detection (SOD) models usually treat RGB and depth as independent information and design separate networks for feature extraction from each. Such schemes can easily be constrained by a limited amount of training data or over-reliance on an elaborately designed training process. Inspired by the observation that RGB and depth modalities actually present certain commonality in distinguishing salient objects, a novel joint learning and densely cooperative fusion (JL-DCF) architecture is designed to learn from both RGB and depth inputs through a shared network backbone, known as the Siamese architecture. In this paper, we propose two effective components: joint learning (JL), and densely cooperative fusion (DCF). The JL module provides robust saliency feature learning by exploiting cross-modal commonality via a Siamese network, while the DCF module is introduced for complementary feature discovery. Comprehensive experiments using five popular metrics show that the designed framework yields a robust RGB-D saliency detector with good generalization. As a result, JL-DCF significantly advances the state-of-the-art models by an average of  2.0 challenging datasets. In addition, we show that JL-DCF is readily applicable to other related multi-modal detection tasks, including RGB-T (thermal infrared) SOD and video SOD (VSOD), achieving comparable or even better performance against state-of-the-art methods. This further confirms that the proposed framework could offer a potential solution for various applications and provide more insight into the cross-modal complementarity task. The code will be available at


page 1

page 4

page 5

page 6

page 9

page 10

page 12

page 13


JL-DCF: Joint Learning and Densely-Cooperative Fusion Framework for RGB-D Salient Object Detection

This paper proposes a novel joint learning and densely-cooperative fusio...

Rethinking RGB-D Salient Object Detection: Models, Datasets, and Large-Scale Benchmarks

The use of RGB-D information for salient object detection has been explo...

Hierarchical Dynamic Filtering Network for RGB-D Salient Object Detection

The main purpose of RGB-D salient object detection (SOD) is how to bette...

Depth Quality-Inspired Feature Manipulation for Efficient RGB-D and Video Salient Object Detection

Recently CNN-based RGB-D salient object detection (SOD) has obtained sig...

Densely Deformable Efficient Salient Object Detection Network

Salient Object Detection (SOD) domain using RGB-D data has lately emerge...

A Unified Multimodal De- and Re-coupling Framework for RGB-D Motion Recognition

Motion recognition is a promising direction in computer vision, but the ...

Depth Quality-Inspired Feature Manipulation for Efficient RGB-D Salient Object Detection

RGB-D salient object detection (SOD) recently has attracted increasing r...

Please sign up or login with your details

Forgot password? Click here to reset