Science Cast

SimPLR: A Simple and Plain Transformer for Object Detection and Segmentation

Duy-Kien NguyenOctober 10, 2023 1:50pm

Views (43)
Comments (0)

Export Citation

Voice is AI-generated

Connected to paperThis paper is a preprint and has not been certified by peer review

SimPLR: A Simple and Plain Transformer for Object Detection and Segmentation

arXivPDFOctober 9, 2023 12:00am

Authors

Duy-Kien Nguyen, Martin R. Oswald, Cees G. M. Snoek

Abstract

The ability to detect objects in images at varying scales has played a pivotal role in the design of modern object detectors. Despite considerable progress in removing handcrafted components using transformers, multi-scale feature maps remain a key factor for their empirical success, even with a plain backbone like the Vision Transformer (ViT). In this paper, we show that this reliance on feature pyramids is unnecessary and a transformer-based detector with scale-aware attention enables the plain detector `SimPLR' whose backbone and detection head both operate on single-scale features. The plain architecture allows SimPLR to effectively take advantages of self-supervised learning and scaling approaches with ViTs, yielding strong performance compared to multi-scale counterparts. We demonstrate through our experiments that when scaling to larger backbones, SimPLR indicates better performance than end-to-end detectors (Mask2Former) and plain-backbone detectors (ViTDet), while consistently being faster. The code will be released.

TwitterandLinkedIn

0 comments

Add comment

SimPLR: A Simple and Plain Transformer for Object Detection and Segmentation

SimPLR: A Simple and Plain Transformer for Object Detection and Segmentation

AI-powered Paper ChatBeta

AI-powered Paper ChatBeta

0 comments