Jahn Heymann scite author profile

This paper introduces a new open source platform for end-toend speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and Py-Torch, as a main deep learning engine. ESPnet also follows the Kaldi ASR toolkit style for data processing, feature extraction/format, and recipes to provide a complete setup for speech recognition and other speech processing experiments. This paper explains a major architecture of this software platform, several important functionalities, which differentiate ESPnet from other open source ASR toolkits, and experimental results with major ASR benchmarks.

show abstract

Neural network based spectral mask estimation for acoustic beamforming

Heymann

2016

View full text Add to dashboard Cite

ESPnet: End-to-End Speech Processing Toolkit

Watanabe

Hori

Karita³

et al. 2018

Preprint

128

134

View full text Add to dashboard Cite

Front-end processing for the CHiME-5 dinner party scenario

Boeddeker¹,

Heitkaemper²,

Schmalenstroeer³

et al. 2018

106

View full text Add to dashboard Cite

This contribution presents a speech enhancement system for the CHiME-5 Dinner Party Scenario. The front-end employs multi-channel linear time-variant filtering and achieves its gains without the use of a neural network. We present an adaptation of blind source separation techniques to the CHiME-5 database which we call Guided Source Separation (GSS). Using the baseline acoustic and language model, the combination of Weighted Prediction Error based dereverberation, guided source separation, and beamforming reduces the WER by 10.54 % (relative) for the single array track and by 21.12 % (relative) on the multiple array track.

show abstract

Beamnet: End-to-end training of a beamformer-supported multi-channel ASR system

et al. 2017

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

customersupport@researchsolutions.com

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Jahn Heymann

ESPnet: End-to-End Speech Processing Toolkit

Neural network based spectral mask estimation for acoustic beamforming

ESPnet: End-to-End Speech Processing Toolkit

Front-end processing for the CHiME-5 dinner party scenario

Beamnet: End-to-end training of a beamformer-supported multi-channel ASR system

Contact Info

Product

Resources

About