Skip to content

add model maxvit and coatnet#3428

Open
learncat163 wants to merge 5 commits into
PaddlePaddle:developfrom
learncat163:rude-maxvit
Open

add model maxvit and coatnet#3428
learncat163 wants to merge 5 commits into
PaddlePaddle:developfrom
learncat163:rude-maxvit

Conversation

@learncat163

Copy link
Copy Markdown

因为 CoAtNet 和maxvit使用的是同一组基础代码,所以这里把相关的权重和介绍说明都放在一起, 把2个模型放在一个PR中提交了。

CoAtNet 模型权重

模型描述

CoAtNet (Coupling Convolution and Attention) 是一种将卷积与 Transformer 注意力相结合的视觉骨干网络。其核心思想是:浅层使用 MBConv 卷积块(利用卷积的平移不变性和局部特征提取能力),深层使用 Transformer 自注意力块(利用其全局建模能力)。CoAtNet 通过这种"卷积先行、注意力殿后"的混合架构,在 ImageNet 分类上取得了优异的性能,同时保持了良好的泛化能力。

模型变体

模型名称 timm 名称 输入尺寸 embed_dim depths rel_pos_type Top-1 (%)
CoAtNet_0_rw_224 coatnet_0_rw_224.sw_in1k 224x224 (152,160,192,224) (2,3,7,2) RelPosBias 78.36
CoAtNet_1_rw_224 coatnet_1_rw_224.sw_in1k 224x224 (152,160,192,224) (2,6,14,2) RelPosBias 81.65
CoAtNet_bn_0_rw_224 coatnet_bn_0_rw_224.sw_in1k 224x224 (152,160,192,224) (2,3,7,2) RelPosBias 82.26
CoAtNet_nano_rw_224 coatnet_nano_rw_224.sw_in1k 224x224 (64,96,128,128) (2,3,5,2) RelPosMlp 76.59
CoAtNet_rmlp_1_rw_224 coatnet_rmlp_1_rw_224.sw_in1k 224x224 (152,160,192,224) (2,6,14,2) RelPosMlp 77.44
CoAtNet_rmlp_2_rw_224 coatnet_rmlp_2_rw_224.sw_in1k 224x224 (192,224,384,512) (2,6,14,2) RelPosMlp 82.93
CoAtNet_rmlp_nano_rw_224 coatnet_rmlp_nano_rw_224.sw_in1k 224x224 (64,96,128,128) (2,3,5,2) RelPosMlp 76.01

预训练权重

权重来源

所有权重均从 timm (HuggingFace) 的 safetensors 格式转换为 PaddlePaddle pdparams 格式。

PaddleClas 模型 timm 模型 权重文件 文件大小
CoAtNet_0_rw_224 coatnet_0_rw_224.sw_in1k coatnet_0_rw_224.sw_in1k.pdparams 105M
CoAtNet_1_rw_224 coatnet_1_rw_224.sw_in1k coatnet_1_rw_224.sw_in1k.pdparams 160M
CoAtNet_bn_0_rw_224 coatnet_bn_0_rw_224.sw_in1k coatnet_bn_0_rw_224.sw_in1k.pdparams 105M
CoAtNet_nano_rw_224 coatnet_nano_rw_224.sw_in1k coatnet_nano_rw_224.sw_in1k.pdparams 58M
CoAtNet_rmlp_1_rw_224 coatnet_rmlp_1_rw_224.sw_in1k coatnet_rmlp_1_rw_224.sw_in1k.pdparams 160M
CoAtNet_rmlp_2_rw_224 coatnet_rmlp_2_rw_224.sw_in1k coatnet_rmlp_2_rw_224.sw_in1k.pdparams 282M
CoAtNet_rmlp_nano_rw_224 coatnet_rmlp_nano_rw_224.sw_in1k coatnet_rmlp_nano_rw_224.sw_in1k.pdparams 58M

使用方式

PaddleClas 中使用

import paddle
from ppcls.arch.backbone.model_zoo.maxxvit import CoAtNet_0_rw_224

# 方式1: 工厂函数直接加载预训练权重
model = CoAtNet_0_rw_224(pretrained=True)

# 方式2: 手动加载转换后的权重
model = CoAtNet_0_rw_224()
state_dict = paddle.load('coatnet_0_rw_224.sw_in1k.pdparams')
model.set_state_dict(state_dict)
model.eval()

x = paddle.randn([1, 3, 224, 224])
output = model(x)  # shape: [1, 1000]

精度对齐

1. 前向推理精度对齐:timm vs Paddle(GPU)

模型 输入尺寸 平均绝对误差
coatnet_0_rw_224.sw_in1k 224x224 5.43e-04
coatnet_1_rw_224.sw_in1k 224x224 3.16e-04
coatnet_bn_0_rw_224.sw_in1k 224x224 5.49e-04
coatnet_nano_rw_224.sw_in1k 224x224 8.10e-05
coatnet_rmlp_1_rw_224.sw_in1k 224x224 1.20e-04
coatnet_rmlp_2_rw_224.sw_in1k 224x224 2.49e-04
coatnet_rmlp_nano_rw_224.sw_in1k 224x224 1.58e-04

2. ImageNet 验证集精度验证

使用5W张 ImageNet进行测试

模型 输入尺寸 Paddle Acc timm Acc 误差
coatnet_0_rw_224.sw_in1k 224x224 78.36% 78.36% 0.01%
coatnet_1_rw_224.sw_in1k 224x224 81.65% 81.65% 0.00%
coatnet_bn_0_rw_224.sw_in1k 224x224 82.26% 82.27% 0.00%
coatnet_nano_rw_224.sw_in1k 224x224 76.59% 76.58% 0.00%
coatnet_rmlp_1_rw_224.sw_in1k 224x224 77.44% 77.44% 0.00%
coatnet_rmlp_2_rw_224.sw_in1k 224x224 82.93% 82.92% 0.00%
coatnet_rmlp_nano_rw_224.sw_in1k 224x224 76.01% 76.02% 0.00%

MaxViT 模型权重

模型描述

MaxViT (Multi-Axis Vision Transformer) 是一种结合卷积与注意力的混合视觉骨干网络。其核心创新是 Multi-Axis Self-Attention (MXSA),将注意力分解为局部的窗口注意力和全局的网格注意力,使模型能够在单层内同时捕获局部和全局信息。MaxViT 在图像分类、目标检测、语义分割等任务上均取得了优异的性能。

模型变体

模型名称 timm 名称 输入尺寸 embed_dim depths stem_width Top-1 (%)
MaxViT_tiny_tf_224 maxvit_tiny_tf_224.in1k 224x224 (64,128,256,512) (2,2,5,2) 64 83.02
MaxViT_tiny_tf_384 maxvit_tiny_tf_384.in1k 384x384 (64,128,256,512) (2,2,5,2) 64 84.20
MaxViT_tiny_tf_512 maxvit_tiny_tf_512.in1k 512x512 (64,128,256,512) (2,2,5,2) 64 84.80
MaxViT_small_tf_224 maxvit_small_tf_224.in1k 224x224 (96,192,384,768) (2,2,5,2) 64 84.00
MaxViT_small_tf_384 maxvit_small_tf_384.in1k 384x384 (96,192,384,768) (2,2,5,2) 64 84.73
MaxViT_small_tf_512 maxvit_small_tf_512.in1k 512x512 (96,192,384,768) (2,2,5,2) 64 85.38
MaxViT_base_tf_224 maxvit_base_tf_224.in1k 224x224 (96,192,384,768) (2,6,14,2) 64 84.69
MaxViT_base_tf_384 maxvit_base_tf_384.in1k 384x384 (96,192,384,768) (2,6,14,2) 64 85.74
MaxViT_base_tf_512 maxvit_base_tf_512.in1k 512x512 (96,192,384,768) (2,6,14,2) 64 86.18
MaxViT_large_tf_224 maxvit_large_tf_224.in1k 224x224 (192,384,768,1536) (2,6,14,2) 128 84.62
MaxViT_large_tf_384 maxvit_large_tf_384.in1k 384x384 (192,384,768,1536) (2,6,14,2) 128 85.53
MaxViT_large_tf_512 maxvit_large_tf_512.in1k 512x512 (192,384,768,1536) (2,6,14,2) 128 85.92

预训练权重

权重来源

所有权重均从 timm (HuggingFace) 的 safetensors 格式转换为 PaddlePaddle pdparams 格式。

PaddleClas 模型 timm 模型 权重文件 文件大小
MaxViT_tiny_tf_224 maxvit_tiny_tf_224.in1k maxvit_tiny_tf_224.in1k.pdparams 119M
MaxViT_tiny_tf_384 maxvit_tiny_tf_384.in1k maxvit_tiny_tf_384.in1k.pdparams 119M
MaxViT_tiny_tf_512 maxvit_tiny_tf_512.in1k maxvit_tiny_tf_512.in1k.pdparams 119M
MaxViT_small_tf_224 maxvit_small_tf_224.in1k maxvit_small_tf_224.in1k.pdparams 264M
MaxViT_small_tf_384 maxvit_small_tf_384.in1k maxvit_small_tf_384.in1k.pdparams 264M
MaxViT_small_tf_512 maxvit_small_tf_512.in1k maxvit_small_tf_512.in1k.pdparams 265M
MaxViT_base_tf_224 maxvit_base_tf_224.in1k maxvit_base_tf_224.in1k.pdparams 457M
MaxViT_base_tf_384 maxvit_base_tf_384.in1k maxvit_base_tf_384.in1k.pdparams 458M
MaxViT_base_tf_512 maxvit_base_tf_512.in1k maxvit_base_tf_512.in1k.pdparams 458M
MaxViT_large_tf_224 maxvit_large_tf_224.in1k maxvit_large_tf_224.in1k.pdparams 809M
MaxViT_large_tf_384 maxvit_large_tf_384.in1k maxvit_large_tf_384.in1k.pdparams 810M
MaxViT_large_tf_512 maxvit_large_tf_512.in1k maxvit_large_tf_512.in1k.pdparams 811M

使用方式

PaddleClas 中使用

import paddle
from ppcls.arch.backbone.model_zoo.maxxvit import MaxViT_tiny_tf_224

# 方式1: 工厂函数直接加载预训练权重
model = MaxViT_tiny_tf_224(pretrained=True)

# 方式2: 手动加载转换后的权重
model = MaxViT_tiny_tf_224()
state_dict = paddle.load('maxvit_tiny_tf_224.in1k.pdparams')
model.set_state_dict(state_dict)
model.eval()

x = paddle.randn([1, 3, 224, 224])
output = model(x)  # shape: [1, 1000]

精度对齐

1. 前向推理精度对齐:timm vs Paddle

模型 输入尺寸 平均绝对误差 Top-1 一致
maxvit_tiny_tf_224.in1k 224x224 5.21e-04 Yes
maxvit_tiny_tf_384.in1k 384x384 2.71e-04 Yes
maxvit_tiny_tf_512.in1k 512x512 9.84e-05 Yes
maxvit_small_tf_224.in1k 224x224 3.11e-04 Yes
maxvit_small_tf_384.in1k 384x384 1.24e-04 Yes
maxvit_small_tf_512.in1k 512x512 1.21e-04 Yes
maxvit_base_tf_224.in1k 224x224 2.71e-04 Yes
maxvit_base_tf_384.in1k 384x384 3.33e-04 Yes
maxvit_base_tf_512.in1k 512x512 4.31e-04 Yes
maxvit_large_tf_224.in1k 224x224 2.08e-04 Yes
maxvit_large_tf_384.in1k 384x384 3.14e-04 Yes
maxvit_large_tf_512.in1k 512x512 1.46e-04 Yes

2. ImageNet 验证集精度验证

模型 输入尺寸 Paddle Acc timm Acc 误差
maxvit_tiny_tf_224.in1k 224x224 83.02% 83.02% 0.00%
maxvit_tiny_tf_384.in1k 384x384 84.20% 84.20% 0.00%
maxvit_tiny_tf_512.in1k 512x512 84.80% 84.80% 0.00%
maxvit_small_tf_224.in1k 224x224 84.00% 84.01% 0.00%
maxvit_small_tf_384.in1k 384x384 84.73% 84.73% 0.00%
maxvit_small_tf_512.in1k 512x512 85.38% 85.38% 0.00%
maxvit_base_tf_224.in1k 224x224 84.69% 84.68% 0.00%
maxvit_base_tf_384.in1k 384x384 85.74% 85.74% 0.00%
maxvit_base_tf_512.in1k 512x512 86.18% 86.19% 0.00%
maxvit_large_tf_224.in1k 224x224 84.62% 84.62% 0.00%
maxvit_large_tf_384.in1k 384x384 85.53% 85.53% 0.00%
maxvit_large_tf_512.in1k 512x512 85.92% 85.92% 0.00%

learncat163 and others added 5 commits June 13, 2026 18:24
Resolve conflict in ppcls/arch/backbone/__init__.py by keeping both
HEAD's maxxvit imports and develop's efficientvit/fastvit/edgenext/
swiftformer imports.
@learncat163

Copy link
Copy Markdown
Author

loss测试

使用 MaxViT_tiny_tf_224 和 CoAtNet_nano_rw_224 测试
数据: lite 训练集 (29 张 ImageNet 样本), batch_size=8, 5 epoch
每 epoch 4 个 batch (最后一批 5 张)
lr 按 batch_size 缩放 (4e-3 for batch4096 -> 1e-4 for batch8)

MaxViT_tiny_tf_224

Epoch 1

Iter LR CELoss
0/4 0.00010000 7.30199
1/4 0.00010000 7.17555
2/4 0.00010000 7.08732
3/4 0.00008200 7.02685
Avg - 7.02685

Epoch 2

Iter LR CELoss
0/4 0.00008200 6.51895
1/4 0.00008200 6.48050
2/4 0.00008200 6.53395
3/4 0.00004789 6.51569
Avg - 6.51569

Epoch 3

Iter LR CELoss
0/4 0.00004789 6.44224
1/4 0.00004789 6.25562
2/4 0.00004789 6.24098
3/4 0.00001905 6.25672
Avg - 6.25672

Epoch 4

Iter LR CELoss
0/4 0.00001905 6.02414
1/4 0.00001905 6.16222
2/4 0.00001905 6.07481
3/4 0.00000361 6.07515
Avg - 6.07515

Epoch 5

Iter LR CELoss
0/4 0.00000361 6.06444
1/4 0.00000361 5.99717
2/4 0.00000361 6.05636
3/4 0.00000100 6.13429
Avg - 6.13429

每 Epoch 平均 Loss

Epoch Avg CELoss LR (起点 -> 末点)
1 7.02685 1.00e-4 -> 8.20e-5
2 6.51569 8.20e-5 -> 4.79e-5
3 6.25672 4.79e-5 -> 1.91e-5
4 6.07515 1.91e-5 -> 3.61e-6
5 6.13429 3.61e-6 -> 1.00e-6

CoAtNet_nano_rw_224

Epoch 1

Iter LR CELoss
0/4 0.00010000 7.21625
1/4 0.00010000 7.26696
2/4 0.00010000 7.23343
3/4 0.00008200 7.25825
Avg - 7.25825

Epoch 2

Iter LR CELoss
0/4 0.00008200 6.83430
1/4 0.00008200 6.65282
2/4 0.00008200 6.73783
3/4 0.00004789 6.75829
Avg - 6.75829

Epoch 3

Iter LR CELoss
0/4 0.00004789 6.86138
1/4 0.00004789 6.54902
2/4 0.00004789 6.36756
3/4 0.00001905 6.38030
Avg - 6.38030

Epoch 4

Iter LR CELoss
0/4 0.00001905 6.01957
1/4 0.00001905 6.11704
2/4 0.00001905 6.10309
3/4 0.00000361 6.11019
Avg - 6.11019

Epoch 5

Iter LR CELoss
0/4 0.00000361 5.81291
1/4 0.00000361 6.02995
2/4 0.00000361 6.00579
3/4 0.00000100 6.00120
Avg - 6.00120

每 Epoch 平均 Loss

Epoch Avg CELoss LR (起点 -> 末点)
1 7.25825 1.00e-4 -> 8.20e-5
2 6.75829 8.20e-5 -> 4.79e-5
3 6.38030 4.79e-5 -> 1.91e-5
4 6.11019 1.91e-5 -> 3.61e-6
5 6.00120 3.61e-6 -> 1.00e-6

@paddle-bot

paddle-bot Bot commented Jul 14, 2026

Copy link
Copy Markdown

Thanks for your contribution!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant