ViewCLR: Learning Self-supervised Video Representation for Unseen Viewpoints

Das, Srijan; Ryoo, Michael S.

doi:10.48550/arxiv.2112.03905

Cited by 1 publication

(2 citation statements)

References 43 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…This pretraining framework is similar to how self-supervision has been benefiting supervised Computer Vision tasks ( [6,8,10,15,31,37,57]): pretrain with self-supervised losses, and then finetune with the downstream task loss. Motivated by them, in this section, we design and benchmark the two-stage pretraining framework, replacing the joint learning framework used in CURL and SAC+AE.…”

Section: Observation On Pretraining Frameworkmentioning

confidence: 99%

“…Several recent work studied such challenges from various directions, including: (1) Inspired by the great success of self-supervised learning (SSL) with images and videos (e.g., [5,6,8,10,14,15,17,21,31,32,37,40,52,54,55,61,71]), some RL methods [1,42,46,59,63,69,81,88] take advantage of self-supervised learning. This is typically done by applying both self-supervised loss and reinforcement learning loss in one batch.…”

Section: Introductionmentioning

confidence: 99%

See 1 more Smart Citation

Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels?

Li¹,

Shang²,

Das³

et al. 2022

Preprint

Self Cite

View full text Add to dashboard Cite

We investigate whether self-supervised learning (SSL) can improve online reinforcement learning (RL) from pixels. We extend the contrastive reinforcement learning framework (e.g., CURL) that jointly optimizes SSL and RL losses and conduct an extensive amount of experiments with various self-supervised losses. Our observations suggest that the existing SSL framework for RL fails to bring meaningful improvement over the baselines only taking advantage of image augmentation when the same amount of data and augmentation is used. We further perform an evolutionary search to find the optimal combination of multiple self-supervised losses for RL, but find that even such a loss combination fails to meaningfully outperform the methods that only utilize carefully designed image augmentations. Often, the use of self-supervised losses under the existing framework lowered RL performances. We evaluate the approach in multiple different environments including a real-world robot environment, and confirm that no single self-supervised loss or image augmentation method can dominate all environments and that the current framework for joint optimization of SSL and RL is limited. Finally, we empirically investigate the pretraining framework for SSL + RL and the properties of representations learned with different approaches.Preprint. Under review.

show abstract