Embodying Pre-Trained Word Embeddings Through Robot Actions

Toyoda, Minori; Suzuki, Kanata; Mori, Hiroki; Hayashi, Yoshikatsu; Ogata, Tetsuya

doi:10.1109/lra.2021.3067862

Cited by 12 publications

(10 citation statements)

References 25 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…During the conversion between symbolic actions/states and texts in a 3D simulator environment [5], pre-training on large non-paired data has shown an improved performance in zero-shot settings. For description-from/to-action translation, [2] realized appropriate action generation that accepts unseen words not included in the dataset. They successfully retrofitted the pretrained word embeddings into multimodal representations incorporating the action modality.…”

Section: A Utilization Of Pre-trained Model In Translation Tasksmentioning

confidence: 99%

“…The present study also adopts this method, following other studies of bidirectional translation between descriptions and actions [1] [2]. A bi-modal autoencoder model has been proposed for the acquisition of multimodal representations of vision and language [11].…”

Section: B Integration Of Multimodal Representationmentioning

confidence: 99%

“…In this study, we use a retrofitted paired recurrent autoencoder (rPRAE [2]) to perform bidirectional translation between descriptions and actions. An overview of the model is presented in Fig.…”

Section: A Model Overviewmentioning

confidence: 99%

“…To achieve this ability, the robot needs to associate real-world objects with linguistic expressions, and complete paired data are generally required to acquire that relationship. In particular, the machine-learning-based approaches [1] [2] require the preparation of a large paired dataset of descriptions and actions. However, collecting and captioning behavioral data is costly, making it difficult to collect large amounts of paired data.…”

Section: Introductionmentioning

confidence: 99%

“…Several joint learning methods for language and motion have been proposed to address these issues. Using contextindependent word embeddings derived from a large corpus, [2] and [3] dealt with unseen words that were not included in the training data. In other studies, stepwise learning methods were proposed to integrate multiple models from different domains [4] [5].…”

Section: Introductionmentioning

confidence: 99%

See 4 more Smart Citations

Learning Bidirectional Translation between Descriptions and Actions with Small Paired Data

Toyoda,

Suzuki,

Hayashi

et al. 2022

Preprint

Self Cite

View full text Add to dashboard Cite

This study achieved bidirectional translation between descriptions and actions using small paired data. The ability to mutually generate descriptions and actions is essential for robots to collaborate with humans in their daily lives. The robot is required to associate real-world objects with linguistic expressions, and large-scale paired data are required for machine learning approaches. However, a paired dataset is expensive to construct and difficult to collect. This study proposes a two-stage training method for bidirectional translation. In the proposed method, we train recurrent autoencoders (RAEs) for descriptions and actions with a large amount of non-paired data. Then, we fine-tune the entire model to bind their intermediate representations using small paired data. Because the data used for pre-training do not require pairing, behavior-only data or a large language corpus can be used. We experimentally evaluated our method using a paired dataset consisting of motion-captured actions and descriptions. The results showed that our method performed well, even when the amount of paired data to train was small. The visualization of the intermediate representations of each RAE showed that similar actions were encoded in a clustered position and the corresponding feature vectors well aligned.

show abstract

Section: A Utilization Of Pre-trained Model In Translation Tasksmentioning

confidence: 99%

Section: B Integration Of Multimodal Representationmentioning

confidence: 99%

Section: A Model Overviewmentioning

confidence: 99%

Section: Introductionmentioning

confidence: 99%