Filling in the gaps: observing gestures conveying additional information can compensate for missing verbal content

Dargue, Nicole; Phillips, Megan; Sweller, Naomi

doi:10.1007/s11251-021-09549-2

Cited by 3 publications

(1 citation statement)

References 54 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Future work should analyze how well this model works with other languages. Second, in addition to the moving lips and faces that provide the visual speech cues, non-verbal gestures and other spontaneous movements (blinking, for example) also have an impact on normal communication ( Cassell et al, 1999 ; Dargue et al, 2021 ; Goldin-Meadow, 1999 ). A model that takes the head movement into account has been developed ( Chen et al, 2020 ) that allows a potential future investigation on what impact of such model can bring in speech comprehension.…”

Section: Discussionmentioning

confidence: 99%

Speech-In-Noise Comprehension is Improved When Viewing a Deep-Neural-Network-Generated Talking Face

Tong

Wenner

et al. 2022

Trends in Hearing

View full text Add to dashboard Cite

Listening in a noisy environment is challenging, but many previous studies have demonstrated that comprehension of speech can be substantially improved by looking at the talker's face. We recently developed a deep neural network (DNN) based system that generates movies of a talking face from speech audio and a single face image. In this study, we aimed to quantify the benefits that such a system can bring to speech comprehension, especially in noise. The target speech audio was masked with signal to noise ratios of −9, −6, −3, and 0 dB and was presented to subjects in three audio-visual (AV) stimulus conditions: (1) synthesized AV: audio with the synthesized talking face movie; (2) natural AV: audio with the original movie from the corpus; and (3) audio-only: audio with a static image of the talker. Subjects were asked to type the sentences they heard in each trial and keyword recognition was quantified for each condition. Overall, performance in the synthesized AV condition fell approximately halfway between the other two conditions, showing a marked improvement over the audio-only control but still falling short of the natural AV condition. Every subject showed some benefit from the synthetic AV stimulus. The results of this study support the idea that a DNN-based model that generates a talking face from speech audio can meaningfully enhance comprehension in noisy environments, and has the potential to be used as a visual hearing aid.

show abstract

Section: Discussionmentioning

confidence: 99%

Speech-In-Noise Comprehension is Improved When Viewing a Deep-Neural-Network-Generated Talking Face

Tong

Wenner

et al. 2022

Trends in Hearing

View full text Add to dashboard Cite

show abstract

Speech-In-Noise Comprehension is Improved When Viewing a Deep-Neural-Network-Generated Talking Face

Tong

Duan

et al. 2022

Preprint

View full text Add to dashboard Cite

Listening in a noisy environment is challenging, but many previous studies have demonstrated that comprehension of speech can be substantially improved by looking at the talker's face. We recently developed a deep neural network (DNN) based system that generates movies of a talking face from speech audio and a single face image. In this study, we aimed to quantify the benefits that such a system can bring to speech comprehension, especially in noise. The target speech audio was masked with signal to noise ratios of -9, -6, -3, and 0 dB and was presented to subjects in three audio-visual (AV) stimulus conditions: 1) synthesized AV: audio with the synthesized talking face movie; 2) natural AV: audio with the original movie from the corpus; and 3) audio-only: audio with a static image of the talker. Subjects were asked to type the sentences they heard in each trial and keyword recognition was quantified for each condition. Overall, performance in the synthesized AV condition fell approximately halfway between the other two conditions, showing a marked improvement over the audio-only control but still falling short of the natural AV condition. Every subject showed some benefit from the synthetic AV stimulus. The results of this study support the idea that a DNN-based model that generates a talking face from speech audio can meaningfully enhance comprehension in noisy environments, and has the potential to be used as a "visual hearing aid."

show abstract

Predictability of Understanding in Explanatory Interactions Based on Multimodal Cues

Türk,

Lazarov,

Wang

et al. 2024

International Conference on Multimodel Interaction

View full text Add to dashboard Cite

Filling in the gaps: observing gestures conveying additional information can compensate for missing verbal content

Cited by 3 publications

References 54 publications

Speech-In-Noise Comprehension is Improved When Viewing a Deep-Neural-Network-Generated Talking Face

Speech-In-Noise Comprehension is Improved When Viewing a Deep-Neural-Network-Generated Talking Face

Speech-In-Noise Comprehension is Improved When Viewing a Deep-Neural-Network-Generated Talking Face

Predictability of Understanding in Explanatory Interactions Based on Multimodal Cues

Contact Info

Product

Resources

About