Vision Transformers, especially Swin Transformer, have become default backbones for various vision tasks but suffer from high memory consumption and training costs. This letter proposes MoR–Swin, a novel architecture that integrates Mixture of Recursions (MoR) into Swin Transformer. An adaptive token-level recursion mechanism dynamically allocates computational depth based on semantic complexity. A recursive window attention module and a lightweight router with load balancing loss are introduced. Extensive experiments on ImageNet classification, COCO detection, and ADE20K segmentation show that MoR–Swin reduces parameters by about 50% and accelerates inference up to twofold at a modest accuracy cost (within about 0.5 points of Swin-B on ImageNet-1K). It provides a new technical pathway for optimizing Vision Transformer models, significantly enhancing their applicability in resource-constrained environments.
更多
查看译文
关键词
Swin Transformer,mixture of recursions,adaptive computation,parameter efficiency