Deformable Attention
Implementation of Deformable Attention in Pytorch from the paper "Vision Transformer with Deformable Attention"
Install / Use
npx skills add lucidrains/deformable-attentionInstalls into whichever agent you are using.
README
<img src="./deformable-attention.png" width="500px"></img>
Deformable Attention
Implementation of Deformable Attention from <a href="https://arxiv.org/abs/2201.00520">this paper</a> in Pytorch, which appears to be an improvement to what was proposed in DETR. The relative positional embedding has also been modified for better extrapolation, using the Continuous Positional Embedding proposed in SwinV2.
Install
$ pip install deformable-attention
Usage
import torch
from deformable_attention import DeformableAttention
attn = DeformableAttention(
dim = 512, # feature dimensions
dim_head = 64, # dimension per head
heads = 8, # attention heads
dropout = 0., # dropout
downsample_factor = 4, # downsample factor (r in paper)
offset_scale = 4, # scale of offset, maximum offset
offset_groups = None, # number of offset groups, should be multiple of heads
offset_kernel_size = 6, # offset kernel size
)
x = torch.randn(1, 512, 64, 64)
attn(x) # (1, 512, 64, 64)
3d deformable attention
import torch
from deformable_attention import DeformableAttention3D
attn = DeformableAttention3D(
dim = 512, # feature dimensions
dim_head = 64, # dimension per head
heads = 8, # attention heads
dropout = 0., # dropout
downsample_factor = (2, 8, 8), # downsample factor (r in paper)
offset_scale = (2, 8, 8), # scale of offset, maximum offset
offset_kernel_size = (4, 10, 10), # offset kernel size
)
x = torch.randn(1, 512, 10, 32, 32) # (batch, dimension, frames, height, width)
attn(x) # (1, 512, 10, 32, 32)
1d deformable attention for good measure
import torch
from deformable_attention import DeformableAttention1D
attn = DeformableAttention1D(
dim = 128,
downsample_factor = 4,
offset_scale = 2,
offset_kernel_size = 6
)
x = torch.randn(1, 128, 512)
attn(x) # (1, 128, 512)
Citation
@misc{xia2022vision,
title = {Vision Transformer with Deformable Attention},
author = {Zhuofan Xia and Xuran Pan and Shiji Song and Li Erran Li and Gao Huang},
year = {2022},
eprint = {2201.00520},
archivePrefix = {arXiv},
primaryClass = {cs.CV}
}
@misc{liu2021swin,
title = {Swin Transformer V2: Scaling Up Capacity and Resolution},
author = {Ze Liu and Han Hu and Yutong Lin and Zhuliang Yao and Zhenda Xie and Yixuan Wei and Jia Ning and Yue Cao and Zheng Zhang and Li Dong and Furu Wei and Baining Guo},
year = {2021},
eprint = {2111.09883},
archivePrefix = {arXiv},
primaryClass = {cs.CV}
}
Related Skills
mcp
Use the `mcp_perplexity-ask_perplexity_search` tools to answer questions. You should use this instead of the `web_search` tool because it is a lot more accurate.
practical-power-systems-synthesis
This skill enables synthesis in the domain of power-systems (engineering). It represents research-level-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform synthesis operations related to power-systems.
semi-supervised-optogenetics-testing
This skill enables testing in the domain of optogenetics (neuroscience). It represents intermediate-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform testing operations related to optogenetics.
data-mining-interpretation-fundamental
This skill enables interpretation in the domain of data-mining (data-science). It represents fundamental-level expertise and is designed for production use in research, industry, and educational contexts. Use this skill when you need to perform interpretation operations related to data-mining.
