Publications

Refine Results

(Filters Applied) Clear All

R&D Areas

R&D Groups

Year

Items per page

Speaker verification using adapted Gaussian mixture models

January 1, 2000

Journal Article

Author:

Douglas A. Reynolds

…

Published in:

Digit. Signal Process., Vol. 10, No. 1-3, January/April/July, 2000, pp. 19-41. (Fifth Annual NIST Speaker Recognition Workshop, 3-4 June 1999.)

Topic:

speaker recognition

R&D area:

Cyber Security and Information Sciences

R&D group:

Artificial Intelligence Technology and Systems

Summary

In this paper we describe the major elements of MIT Lincoln Laboratory's Gaussian mixture model (GMM)-based speaker verification system used successfully in several NIST Speaker Recognition Evaluations (SREs). The system is built around the likelihood ratio test for verification, using simple but effective GMMs for likelihood functions, a universal background model (UBM) for alternative speaker representation, and a form of Bayesian adaptation to derive speaker models from the UBM. The development and use of a handset detector and score normalization to greatly improve verification performance is also described and discussed. Finally, representative performance benchmarks and system behavior experiments on NIST SRE corpora are presented.

READ LESS

Summary

Speaker verification using adapted Gaussian mixture models

Estimation of modulation based on FM-to-AM transduction: two-sinusoid case

November 1, 1999

Journal Article

Author:

Wade P. Torres

…

Thomas F. Quatieri

Published in:

IEEE Trans. Signal Process., Vol. 47, No. 11, November 1999, pp. 3084-3097.

Topic:

signal processing

R&D area:

Cyber Security and Information Sciences

R&D group:

Artificial Intelligence Technology and Systems

Summary

A method is described for estimating the amplitude modulation (AM) and the frequency modulation (FM) of the components of a signal that consists of two AM-FM sinusoids. The approach is based on the transduction of FM to AM that occurs whenever a signal of varying frequency passes through a filter with a nonflat frequency response. The objective is to separate the AM and FM of the sinusoids from the amplitude envelopes of the output of two transduction filters, where the AM and FM are nonlinearly combined in the amplitude envelopes. A current scheme is first refined for AM-FM estimation of a single AM-FM sinusoid by iteratively inverting the AM and FM estimates to reduce error introduced in transduction. The transduction filter pair is designed relying on both a time-and frequency-domain characterization of transduction error. The approach is then extended to the case of two AM-FM sinusoids by essentially reducing the problem to two single-component AM-FM estimation problems. By exploiting the beating in the amplitude envelope of each filter output due to the two-sinusoidal input, a closed-form solution is obtained. This solution is also improved upon by iterative refinement. The AM-FM estimation methods are evaluated through an error analysis and are illustrated for a wide range of AM-FM signals.

READ LESS

Summary

Estimation of modulation based on FM-to-AM transduction: two-sinusoid case

Shunting networks for multi-band AM-FM decomposition

October 17, 1999

Conference Paper

Author:

Robert A. Baxter

…

Thomas F. Quatieri

Published in:

Proc. IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, 17-20 October 1999.

Topic:

signal processing

R&D area:

Cyber Security and Information Sciences

R&D group:

Artificial Intelligence Technology and Systems

Summary

We describe a transduction-based, neurodynamic approach to estimating the amplitude-modulated (AM) and frequency-modulated (FM) components of a signal. We show that the transduction approach can be realized as a bank of constant-Q bandpass filters followed by envelope detectors and shunting neural networks, and the resulting dynamical system is capable of robust AM-FM estimation. Our model is consistent with recent psychophysical experiments that indicate AM and FM components of acoustic signals may be transformed into a common neural code in the brain stem via FM-to-AM transduction. The shunting network for AM-FM decomposition is followed by a contrast enhancement shunting network that provides a mechanism for robustly selecting auditory filter channels as the FM of an input stimulus sweeps across the multiple filters. The AM-FM output of the shunting networks may provide a robust feature representation and is being considered for applications in signal recognition and multi-component decomposition problems.

READ LESS

Summary

Shunting networks for multi-band AM-FM decomposition

A study of computation speed-ups of the GMM-UBM speaker recognition system

September 5, 1999

Conference Paper

Author:

John J. McLaughlin

…

Published in:

6th European Conf. on Speech Communication and Technology, EUROSPEECH, 5-9 September 1999.

Topic:

speaker recognition

R&D area:

Cyber Security and Information Sciences

R&D group:

Artificial Intelligence Technology and Systems

Summary

The Gaussian Mixture Model Universal Background Model (GMM-UBM) speaker recognition system has demonstrated very high performance in several NIST evaluations. Such evaluations, however, are concerned only with classification accuracy. In many applications, system effectiveness must be evaluated in light of both accuracy and execution speed. We present here a number of techniques for decreasing computation. Using data from the Switchboard telephone speech corpus, we show that significant speed-ups can be obtained while sacrificing surprisingly little accuracy. We expect that these techniques, involving lowering model order as well as processing fewer speech frames, will apply equally well to other recognition systems.

READ LESS

Summary

A study of computation speed-ups of the GMM-UBM speaker recognition system

Evaluation of confidence measures for language identification

September 5, 1999

Conference Paper

Author:

Kay M. Berkling

…

Published in:

6th European Conf. on Speech Communication and Technology, EUROSPEECH, 5-9 September 1999.

Topic:

language recognition

R&D area:

Cyber Security and Information Sciences

R&D group:

Artificial Intelligence Technology and Systems

Summary

In this paper we examine various ways to derive confidence measures for a language identification system, using phone recognition followed by language models, and describe the application of an evaluation metric for measuring the "goodness" of the different confidence measures. Experiments are conducted on the 1996 NIST Language Identification Evaluation corpus (derived from the Callfriend corpus of conversational telephone speech). The system is trained on the NIST 96 development data and evaluated on the NIST 96 evaluation data. Results indicate that we are able to predict the performance of a system and quantitatively evaluate how well the prediction holds on new data.

READ LESS

Summary

Evaluation of confidence measures for language identification

Speaker and language recognition using speech codec parameters

September 5, 1999

Conference Paper

Author:

Thomas F. Quatieri

…

Published in:

EUROSPEECH 99, 5-10 September 1999.

Topic:

speaker recognition

R&D area:

Cyber Security and Information Sciences

R&D group:

Artificial Intelligence Technology and Systems

Summary

In this paper, we investigate the effect of speech coding on speaker and language recognition tasks. Three coders were selected to cover a wide range of quality and bit rates: GSM at 12.2 kb/s, G.729 at 8 kb/s, and G.723.1 at 5.3 kb/s. Our objective is to measure recognition performance from either the synthesized speech or directly from the coder parameters themselves. We show that using speech synthesized from the three codecs, GMM-based speaker verification and phone-based language recognition performance generally degrades with coder bit rate, i.e., from GSM to G.729 to G.723.1, relative to an uncoded baseline. In addition, speaker verification for all codecs shows a performance decrease as the degree of mismatch between training and testing conditions increases, while language recognition exhibited no decrease in performance. We also present initial results in determining the relative importance of codec system components in their direct use for recognition tasks. For the G.729 codec, it is shown that removal of the post- filter in the decoder helps speaker verification performance under the mismatched condition. On the other hand, with use of G.729 LSF-based mel-cepstra, performance decreases under all conditions, indicating the need for a residual contribution to the feature representation.

READ LESS

Summary

Speaker and language recognition using speech codec parameters

Modeling of the glottal flow derivative waveform with application to speaker identification

September 1, 1999

Journal Article

Author:

M. D. Plumpe

…

Published in:

IEEE Trans. Speech Audio Process., Vol. 7, No. 5, September 1999, pp. 569-586.

Topic:

speaker recognition

R&D area:

Cyber Security and Information Sciences

R&D group:

Artificial Intelligence Technology and Systems

Summary

An automatic technique for estimating and modeling the glottal flow derivative source waveform from speech, and applying the model parameters to speaker identification, is presented. The estimate of the glottal flow derivative is decomposed into coarse structure, representing the general flow shape, and fine structure, comprising aspiration and other perturbations in the flow, from which model parameters are obtained. The glottal flow derivative is estimated using an inverse filter determined within a time interval of vocal-fold closure that is identified through differences in formant frequency modulation during the open and closed phases of the glottal cycle. This formant motion is predicted by Ananthapadmanabha and Fant to be a result of time-varying and nonlinear source/vocal tract coupling within a glottal cycle. The glottal flow derivative estimate is modeled using the Liljencrants-Fant model to capture its coarse structure, while the fine structure of the flow derivative is represented through energy and perturbation measures. The model parameters are used in a Gaussian mixture model speaker identification (SID) system. Both coarse- and fine-structure glottal features are shown to contain significant speaker-dependent information. For a large TIMIT database subset, averaging over male and female SID scores, the coarse-structure parameters achieve about 60% accuracy, the fine-structure parameters give about 40% accuracy, and their combination yields about 70% correct identification. Finally, in preliminary experiments on the counterpart telephone-degraded NTIMIT database, about a 5% error reduction in SID scores is obtained when source features are combined with traditional mel-cepstral measures.

READ LESS

Summary

Modeling of the glottal flow derivative waveform with application to speaker identification

Understanding-based translingual information retrieval

June 17, 1999

Conference Paper

Author:

Young-Suk Lee

…

Published in:

4th Int. Conf. on Applications of Natural Language to Information Systems, 17-19 June 1999, pp. 187-195.

Topic:

machine translation

R&D area:

Cyber Security and Information Sciences

R&D group:

Artificial Intelligence Technology and Systems

Summary

This paper describes our preliminary research on an understanding-based translingual information retrieval system for which the input to the system is a query sentence in English, and the output of the system is a set of documents either in English or in Korean. The understanding module produces a meaning representation --- called semantic frame --- of the input sentence where the predicate-argument structure and the question-type of the input are identified, and each keyword is assigned its concept category. The translingual search module performs search on an English and Korean bilingual corpus tagged with concept categories. The results of our preliminary experiment, performed an a document set consisting of slides and notes from English and Korean briefings in a military domain, indicate that an understanding-based approach to information retrieval combined with concept-based search technique improves both precision and recall compared with a keyword match technique without understanding for both monolingual- and translingual retrieval. Current work is directed at further development of the system, and in preparation for tests on larger copora.

READ LESS

Summary

Understanding-based translingual information retrieval

Security implications of adaptive multimedia distribution

June 6, 1999

Conference Paper

Author:

Thomas M. Parks

…

Published in:

Proc. IEEE Int. Conf. on Communications, Multimedia and Wireless, Vol. 3, 6-10 June 1999, pp. 1563-1567.

Topic:

communications

R&D area:

Cyber Security and Information Sciences

R&D group:

Artificial Intelligence Technology and Systems

Summary

We discuss the security implications of different techniques used in adaptive audio and video distribution. Several sources of variability in the network make it necessary for applications to adapt. Ideally, each receiver should receive media quality commensurate with the capacity of the path leading to it from each sender. Several different techniques have been proposed to provide such adaptation. We discuss the implications of each technique for confidentiality, authentication, integrity, and anonymity. By coincidence, the techniques with better performance also have better security properties.

READ LESS

Summary

Security implications of adaptive multimedia distribution

Corpora for the evaluation of speaker recognition systems

March 15, 1999

Conference Paper

Author:

Joseph P. Campbell Jr

…

Douglas A. Reynolds

Published in:

ICASSP 1999, Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing, 15-19 March 1999.

Topic:

human language technology

R&D area:

Cyber Security and Information Sciences

R&D group:

Artificial Intelligence Technology and Systems

Summary

Using standard speech corpora for development and evaluation has proven to be very valuable in promoting progress in speech and speaker recognition research. In this paper, we present an overview of current publicly available corpora intended for speaker recognition research and evaluation. We outline the corpora's salient features with respect to their suitability for conducting speaker recognition experiments and evaluations. Links to these corpora, and to new corpora, will appear on the web http://www.apl.jhu.edu/Classes/Notes/Campbell/SpkrRec/. We hope to increase the awareness and use of these standard corpora and corresponding evaluation procedures throughout the speaker recognition community.

READ LESS

Summary

Corpora for the evaluation of speaker recognition systems

Publications

Refine Results

Summary

Summary

Summary

Summary

Summary

Summary

Summary

Summary

Summary

Summary

Summary

Summary

Summary

Summary

Understanding-based translingual information retrieval

Summary

Summary

Summary

Summary

Summary

Summary

Showing Results