Revisiting the Critical Factors of Augmentation-Invariant Representation Learning

Huang, Junqiang; Kong, Xiangwen; Zhang, Xiangyu

doi:10.1007/978-3-031-19821-2_3

Cited by 3 publications

(2 citation statements)

References 23 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…It has been a notorious problem that fine-tuning LARStrained models with optimization hyper-parameters best selected for SGD-trained counterparts yields sub-optimal performance [11,31]. To address this, recent work [31] proposes NormRescale to scale the norm of LARS-trained weights by a specific anchor (e.g., the norm of SGD-trained weights or a constant number). It helps the LARS-trained models fit to the optimization strategy of fine-tuning.…”

Section: Evaluation Setupmentioning

confidence: 99%

See 1 more Smart Citation

Pixel-Wise Contrastive Distillation

Huang¹,

Guo²

2022

Preprint

View full text Add to dashboard Cite

We present the first pixel-level self-supervised distillation framework specified for dense prediction tasks. Our approach, called Pixel-Wise Contrastive Distillation (PCD), distills knowledge by attracting the corresponding pixels from student's and teacher's output feature maps. This pixel-to-pixel distillation demands for maintaining the spatial information of teacher's output. We propose a Spatial-Adaptor that adapts the well-trained projection/prediction head of the teacher used to encode vectorized features to processing 2D feature maps. SpatialAdaptor enables more informative pixel-level distillation, yielding a better student for dense prediction tasks. Besides, in light of the inadequate effective receptive fields of small models, we utilize a plug-in multi-head self-attention module to explicitly relate the pixels of student's feature maps. Overall, our PCD outperforms previous self-supervised distillation methods on various dense prediction tasks. A backbone of ResNet-18 distilled by PCD achieves 37.4 AP bbox and 34.0 AP mask with Mask R-CNN detector on COCO dataset, emerging as the first pre-training method surpassing the supervised pretrained counterpart.

show abstract

Section: Evaluation Setupmentioning

confidence: 99%

“…Most of checkpoints are from the official repository, except for BYOL [21]. The ResNet-50 [25] pre-trained by BYOL are from the implementation of [31]. We list out the URLs for downloading these models:…”

Section: A3 Distilling From Different Teachersmentioning

confidence: 99%