Daikun Liu scite author profile

Ground-to-aerial geolocalization refers to localizing a ground-level query image by matching it to a reference database of geo-tagged aerial imagery. This is very challenging due to the huge perspective differences in visual appearances and geometric configurations between these two views. In this work, we propose a novel Transformer-guided convolutional neural network (TransGCNN) architecture, which couples CNN-based local features with Transformer-based global representations for enhanced representation learning. Specifically, our TransGCNN consists of a CNN backbone extracting feature map from an input image and a Transformer head modeling global context from the CNN map. In particular, our Transformer head acts as a spatialaware importance generator to select salient CNN features as the final feature representation. Such a coupling procedure allows us to leverage a lightweight Transformer network to greatly enhance the discriminative capability of the embedded features. Furthermore, we design a dual-branch Transformer head network to combine image features from multi-scale windows in order to improve details of the global feature representation. Extensive experiments on popular benchmark datasets demonstrate that our model achieves top-1 accuracy of 94.12% and 84.92% on CVUSA and CVACT val, respectively, which outperforms the second-performing baseline with less than 50% parameters and almost 2× higher frame rate, therefore achieving a preferable accuracy-efficiency tradeoff.

show abstract

Voxel-Based Multi-Scale Transformer Network for Event Stream Processing

Liu

Wang

Sun

2024

IEEE Trans. Circuits Syst. Video Technol.

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

customersupport@researchsolutions.com

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Daikun Liu

Adsorption of SF6 decomposition components over Pd (1 1 1): A density functional theory study

DFT-based study on H2S and SOF2 adsorption on Si-MoS2 monolayer

Transformer-Guided Convolutional Neural Network for Cross-View Geolocalization

Voxel-Based Multi-Scale Transformer Network for Event Stream Processing

Contact Info

Product

Resources

About

Daikun Liu

Adsorption of SF6 decomposition components over Pd (1 1 1): A density functional theory study

DFT-based study on H2S and SOF2 adsorption on Si-MoS2 monolayer

Transformer-Guided Convolutional Neural Network for Cross-View Geolocalization

Voxel-Based Multi-Scale Transformer Network for Event Stream Processing

Contact Info

Product

Resources

About

Adsorption of SF6 decomposition components over Pd (1 1 1): A density functional theory study