Language-Driven Robot Manipulation With Perspective Disambiguation and Placement Optimization

Huang, Kaixiang; Han, Yanshuo; Wu, J.C.; Qiu, Fangzhou; Tang, Qiulin

doi:10.1109/lra.2022.3146955

Cited by 7 publications

(1 citation statement)

References 23 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Additionally, deep learning approaches typically require offline datasets comprising language instructions and visual data (such as images or videos). Existing works Hristov et al (2017) , Ahn et al (2018) , Hatori et al (2018) , Magassouba et al (2019) , Chen et al (2020) , Mi et al (2020a) , Shridhar, Mittal & Hsu (2020) , Mi et al (2020b) , Nguyen et al (2020) , Shridhar et al (2020) , Huang et al (2022) have created various datasets under differing conditions for model training and testing. Although synthetic datasets are more accessible in a simulated environment, methods trained on synthetic data must address the challenge of domain shift problem when applied in a real-world scenarios.…”

Section: Introductionmentioning

confidence: 99%

Knowledge enhanced bottom-up affordance grounding for robotic interaction

Qu,

Li,

Jin

2024

PeerJ Computer Science

View full text Add to dashboard Cite

With the rapid advancement of robotics technology, an increasing number of researchers are exploring the use of natural language as a communication channel between humans and robots. In scenarios where language conditioned manipulation grounding, prevailing methods rely heavily on supervised multimodal deep learning. In this paradigm, robots assimilate knowledge from both language instructions and visual input. However, these approaches lack external knowledge for comprehending natural language instructions and are hindered by the substantial demand for a large amount of paired data, where vision and language are usually linked through manual annotation for the creation of realistic datasets. To address the above problems, we propose the knowledge enhanced bottom-up affordance grounding network (KBAG-Net), which enhances natural language understanding through external knowledge, improving accuracy in object grasping affordance segmentation. In addition, we introduce a semi-automatic data generation method aimed at facilitating the quick establishment of the language following manipulation grounding dataset. The experimental results on two standard dataset demonstrate that our method outperforms existing methods with the external knowledge. Specifically, our method outperforms the two-stage method by 12.98% and 1.22% of mIoU on the two dataset, respectively. For broader community engagement, we will make the semi-automatic data construction method publicly available at https://github.com/wmqu/Automated-Dataset-Construction4LGM.

show abstract