R&D Result Detail

Original Title

SpeakerBeam: Speaker Aware Neural Network for Target Speaker Extraction in Speech Mixtures

English Title

SpeakerBeam: Speaker Aware Neural Network for Target Speaker Extraction in Speech Mixtures

Type

WoS Article

Original Abstract

The processing of speech corrupted by interferingoverlapping speakers is one of the challenging problems withregards to todays automatic speech recognition systems. Recently,approaches based on deep learning have made great progresstoward solving this problem. Most of these approaches tacklethe problem as speech separation, i.e., they blindly recover allthe speakers from the mixture. In some scenarios, such as smartpersonal devices, we may however be interested in recovering onetarget speaker froma mixture. In this paper, we introduce Speaker-Beam, a method for extracting a target speaker from the mixturebased on an adaptation utterance spoken by the target speaker.Formulating the problem as speaker extraction avoids certainissues such as label permutation and the need to determine thenumber of speakers in the mixture.With SpeakerBeam, we jointlylearn to extract a representation from the adaptation utterancecharacterizing the target speaker and to use this representationto extract the speaker. We explore several ways to do this, mostlyinspired by speaker adaptation in acoustic models for automaticspeech recognition. We evaluate the performance on the widelyused WSJ0-2mix andWSJ0-3mix datasets, and these datasets modifiedwith more noise or more realistic overlapping patterns. Wefurther analyze the learned behavior by exploring the speaker representationsand assessing the effect of the length of the adaptationdata. The results show the benefit of including speaker informationin the processing and the effectiveness of the proposed method.

English abstract

The processing of speech corrupted by interferingoverlapping speakers is one of the challenging problems withregards to todays automatic speech recognition systems. Recently,approaches based on deep learning have made great progresstoward solving this problem. Most of these approaches tacklethe problem as speech separation, i.e., they blindly recover allthe speakers from the mixture. In some scenarios, such as smartpersonal devices, we may however be interested in recovering onetarget speaker froma mixture. In this paper, we introduce Speaker-Beam, a method for extracting a target speaker from the mixturebased on an adaptation utterance spoken by the target speaker.Formulating the problem as speaker extraction avoids certainissues such as label permutation and the need to determine thenumber of speakers in the mixture.With SpeakerBeam, we jointlylearn to extract a representation from the adaptation utterancecharacterizing the target speaker and to use this representationto extract the speaker. We explore several ways to do this, mostlyinspired by speaker adaptation in acoustic models for automaticspeech recognition. We evaluate the performance on the widelyused WSJ0-2mix andWSJ0-3mix datasets, and these datasets modifiedwith more noise or more realistic overlapping patterns. Wefurther analyze the learned behavior by exploring the speaker representationsand assessing the effect of the length of the adaptationdata. The results show the benefit of including speaker informationin the processing and the effectiveness of the proposed method.

Keywords

Speaker extraction, speaker-aware neural network,multi-speaker speech recognition.

Key words in English

Speaker extraction, speaker-aware neural network,multi-speaker speech recognition.

Authors

ŽMOLÍKOVÁ, K.; DELCROIX, M.; KINOSHITA, K.; OCHIAI, T.; NAKATANI, T.; BURGET, L.; ČERNOCKÝ, J.

RIV year

2020

Released

13.06.2019

ISBN

1932-4553

Periodical

IEEE Journal of Selected Topics in Signal Processing

Volume

13

Number

4

State

United States of America

Pages from

800

Pages to

814

Pages count

15

URL

https://ieeexplore.ieee.org/document/8736286

BibTex

@article{BUT159990,
  author="ŽMOLÍKOVÁ, K. and DELCROIX, M. and KINOSHITA, K. and OCHIAI, T. and NAKATANI, T. and BURGET, L. and ČERNOCKÝ, J.",
  title="SpeakerBeam: Speaker Aware Neural Network for Target Speaker Extraction in Speech Mixtures",
  journal="IEEE Journal of Selected Topics in Signal Processing",
  year="2019",
  volume="13",
  number="4",
  pages="800--814",
  doi="10.1109/JSTSP.2019.2922820",
  issn="1932-4553",
  url="https://ieeexplore.ieee.org/document/8736286"
}

Documents

zmolikova_IEEEjournal2019_08736286
SpeakerBeam

VUT

Faculties and university institutes

Parts

SpeakerBeam: Speaker Aware Neural Network for Target Speaker Extraction in Speech Mixtures