Wave-ViT: Unifying wavelet and transformers for visual representation learning
Multi-scale Vision Transformer (ViT) has emerged as a powerful backbone for computer vision tasks, while the self-attention computation in Transformer scales quadratically w.r.t. the input patch number. Thus, existing solutions commonly employ down-sampling operations (e.g., average pooling) over ke...
Saved in:
Main Authors: | YAO, Ting, PAN, Yingwei, LI, Yehao, NGO, Chong-wah, MEI, Tao |
---|---|
格式: | text |
語言: | English |
出版: |
Institutional Knowledge at Singapore Management University
2022
|
主題: | |
在線閱讀: | https://ink.library.smu.edu.sg/sis_research/7508 https://ink.library.smu.edu.sg/context/sis_research/article/8511/viewcontent/2207.04978.pdf |
標簽: |
添加標簽
沒有標簽, 成為第一個標記此記錄!
|
機構: | Singapore Management University |
語言: | English |
相似書籍
-
Face recognition by applying wavelet subband representation and kernel associative memory
由: Zhang, B.-L., et al.
出版: (2014) -
The application of wavelet transform to analyze the rainfall data
由: Nareemal Hilae
出版: (2018) -
Exponential B-Splines: Scale-Space and Wavelet Representations
由: LOR CHOON YEE
出版: (2012) -
Unifying global-local representations in salient object detection with transformers
由: REN, Sucheng, et al.
出版: (2024) -
Efficient architecture for discrete wavelet transform using daubechies
由: Zhang Xiaoyin
出版: (2011)