ScaleFormer: Revisiting the Transformer-based Backbones from a Scale-wise Perspective for Medical Image Segmentation

by   Huimin Huang, et al.

Recently, a variety of vision transformers have been developed as their capability of modeling long-range dependency. In current transformer-based backbones for medical image segmentation, convolutional layers were replaced with pure transformers, or transformers were added to the deepest encoder to learn global context. However, there are mainly two challenges in a scale-wise perspective: (1) intra-scale problem: the existing methods lacked in extracting local-global cues in each scale, which may impact the signal propagation of small objects; (2) inter-scale problem: the existing methods failed to explore distinctive information from multiple scales, which may hinder the representation learning from objects with widely variable size, shape and location. To address these limitations, we propose a novel backbone, namely ScaleFormer, with two appealing designs: (1) A scale-wise intra-scale transformer is designed to couple the CNN-based local features with the transformer-based global cues in each scale, where the row-wise and column-wise global dependencies can be extracted by a lightweight Dual-Axis MSA. (2) A simple and effective spatial-aware inter-scale transformer is designed to interact among consensual regions in multiple scales, which can highlight the cross-scale dependency and resolve the complex scale variations. Experimental results on different benchmarks demonstrate that our Scale-Former outperforms the current state-of-the-art methods. The code is publicly available at:


page 2

page 3

page 4

page 6

page 7


C2FTrans: Coarse-to-Fine Transformers for Medical Image Segmentation

Convolutional neural networks (CNN), the most prevailing architecture fo...

AerialFormer: Multi-resolution Transformer for Aerial Image Segmentation

Aerial Image Segmentation is a top-down perspective semantic segmentatio...

Dual-Flattening Transformers through Decomposed Row and Column Queries for Semantic Segmentation

It is critical to obtain high resolution features with long range depend...

MISSFormer: An Effective Medical Image Segmentation Transformer

The CNN-based methods have achieved impressive results in medical image ...

Enhancing Medical Image Segmentation with TransCeption: A Multi-Scale Feature Fusion Approach

While CNN-based methods have been the cornerstone of medical image segme...

Class-Aware Generative Adversarial Transformers for Medical Image Segmentation

Transformers have made remarkable progress towards modeling long-range d...

SwinVFTR: A Novel Volumetric Feature-learning Transformer for 3D OCT Fluid Segmentation

Accurately segmenting fluid in 3D volumetric optical coherence tomograph...

Please sign up or login with your details

Forgot password? Click here to reset