Unsupervised Learning of a Disentangled Speech Representation for Voice Conversion

Gburrek, Tobias; Glarner, Thomas; Ebbers, Janek; Haeb‐Umbach, Reinhold; Wagner, Petra

doi:10.21437/ssw.2019-15

Cited by 4 publications

(7 citation statements)

References 12 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Among the different models rising in the field in deep generative modeling, variational auto-encoder [17,18] has become probably the most popular in speech processing [19,20]. Variational auto-encoder is a deep latent variable model, where the probability density of the observation depends on the output of a highly non-linear transformation of the latent variable, where the transformation is modeled by a neural network (known as a decoder).…”

Section: Variational Autoencoder As Spectral Modelmentioning

confidence: 99%

“…In our work, we use VAE to model single-speaker speech in the log-magnitude STFT domain. We use a (slightly modified) model from [20], which was inspired by [19]. In this model, part of the latent variables is trained to model speaker variability.…”

Section: Variational Autoencoder As Spectral Modelmentioning

confidence: 99%

“…Note that the speaker classification objective differs from [19,20], where the aim was to learn to disentangle the speaker information in an unsupervised way. Here, we assume having the speaker identities during training, and we can thus directly use a supervised objective.…”

Section: Variational Autoencoder As Spectral Modelmentioning

confidence: 99%

“…The value inside the expectation expresses how likely the source k is dominant and has the value of the observed mixed speech y t,f . For example of k = 1, this can be expressed as ln p(y t,f , d t,f = 1|Z (1) , Z (2) (20) where Φ(•) is the cumulative normal distribution. This solution can be also found in previous literature on lifted-max model [7,11].…”

Section: Inference Of Q(d)mentioning

confidence: 99%

“…To combine the strengths of both stated directions, we here propose to use a neural network as a generative model of the speech signal in a factorial model. For this, we use the variational autoencoder (VAE) [17,18], which has been previously successfully used to model speech [19,20]. We train VAE on clean single-speaker speech and create a simple Gaussian model of noise.…”

Section: Introductionmentioning

confidence: 99%

See 4 more Smart Citations

Integration of variational autoencoder and spatial clustering for adaptive multi-channel neural speech separation

Žmolíková¹,

Delcroix²,

Burget³

et al. 2020

Preprint

View full text Add to dashboard Cite

In this paper, we propose a method combining variational autoencoder model of speech with a spatial clustering approach for multichannel speech separation. The advantage of integrating spatial clustering with a spectral model was shown in several works. As the spectral model, previous works used either factorial generative models of the mixed speech or discriminative neural networks. In our work, we combine the strengths of both approaches, by building a factorial model based on a generative neural network, a variational autoencoder. By doing so, we can exploit the modeling power of neural networks, but at the same time, keep a structured model. Such a model can be advantageous when adapting to new noise conditions as only the noise part of the model needs to be modified. We show experimentally, that our model significantly outperforms previous factorial model based on Gaussian mixture model (DOLPHIN), performs comparably to integration of permutation invariant training with spatial clustering, and enables us to easily adapt to new noise conditions.

show abstract

Section: Variational Autoencoder As Spectral Modelmentioning

confidence: 99%

Section: Variational Autoencoder As Spectral Modelmentioning

confidence: 99%

Section: Variational Autoencoder As Spectral Modelmentioning

confidence: 99%