Developing of a Software–Hardware Complex for Automatic Audio–Visual Speech Recognition in Human–Robot Interfaces

In this paper we present a novel method designed for word-level visual speech recognition and intended for use in human-robot interaction. The ability of robots to understand natural human speech will significantly improve the quality of human-machine interaction. Despite outstanding breakthroughs achieved in this field in recent years this challenge remains unresolved. In current research we mainly focus on the visual part of the human speech, so-called automated lip-reading task, which becomes crucial for human-robot interaction in acoustically noisy environment. The developed method is based on the use of state-of-the-art artificial intelligence technologies and allowed to achieve an incredible 85.03% speech recognition accuracy using only video data. It is worth noting that the model training and testing of the method was carried out on a benchmarking LRW database recorded inthe-wild, and the presented results surpass many existing achieved by the researchers of the world speech recognition community.

show abstract

Developing of a Software–Hardware Complex for Automatic Audio–Visual Speech Recognition in Human–Robot Interfaces

Cited by 3 publications

References 22 publications

Analyzing lower half facial gestures for lip reading applications: Survey on vision techniques

Analyzing lower half facial gestures for lip reading applications: Survey on vision techniques

End-to-end visual speech recognition for human-robot interaction

End-to-end Visual Speech Recognition for Human-Robot Interaction

Contact Info

Product

Resources

About