Nonstationary time series modeling with applications to speech signal processing

Nonstationary time series modeling with applications to speech signal processing

by Daniel Rudoy

Part of Collections of the Harvard University Archives

Browse books you can read free on Readfeed

No club is reading this yet — be the first to start one

Start a club free
About
We develop statistical methods for the analysis of nonstationary time series and apply them to a variety of problems arising in speech signal processing. Information-carrying natural sound signals such as speech exhibit a degree of controlled nonstationarity in that their statistical properties vary slowly over time. Faithfully modeling these temporal variations is extremely valuable for a wide range of applications and can be accomplished by relying on well-understood acoustic models of speech production, which motivate many of the methods developed in this thesis. First, we make a number of contributions to the classical problem of formant tracking, in which vocal tract resonances are estimated under the assumption of their invariance on the 15-30 ms scale. Next, we relax this piecewise-stationarity constraint and model the temporal dynamics of the vocal tract using time-varying autoregressive (TVAR) models. We develop their algebraic and geometric properties, introduce several new estimators, and use TVAR models to develop a hypothesis test to detect the presence of vocal tract variation in speech waveform data. We study its asymptotic properties, and illustrate its practical efficacy by detecting vocal tract changes across different timescales of speech dynamics. Next, we explore how standard fixed-resolution short-time Fourier representations may be generalized in order to adapt to the time-frequency structure of a speech signal. To this end, we introduce a family of adaptive, linear time-frequency representations termed superposition frames and show that they are invertible, numerically-stable, and admit fast overlap-add reconstruction akin to standard short-time Fourier techniques. The general construction proceeds via a local signal-adaptive modification of a Gabor frame. Two signal-dependent schemes for selecting an appropriate superposition frame for signal analysis are given, and the framework is illustrated in the context of speech enhancement. Finally, we introduce a joint model of the vocal tract and the source waveform in order to take into account its quasi-periodic temporal variations during voicing. We incorporate an estimate of the source waveform into the traditional linear prediction framework via nonparametric wavelet regression; the resultant semi-parametric model is applied to various speech analysis problems including formant and source-harmonics-to-noise ratio estimation, inverse filtering, and voicing detection.

Discuss Nonstationary time series modeling with applications to speech signal processing with other readers

Join or start a book club for Nonstationary time series modeling with applications to speech signal processing on Readfeed. Live chat, shared reading progress, and AI discussion questions — free to get started.

Frequently asked questions

How do I join a book club for Nonstationary time series modeling with applications to speech signal processing?

Sign up free on Readfeed, then browse public clubs or start your own club with Nonstationary time series modeling with applications to speech signal processing as the current read. Invite friends with a share link and discuss together with live chat and AI discussion questions.

Can I discuss Nonstationary time series modeling with applications to speech signal processing with other readers online?

Yes. Readfeed book clubs let you chat live, share progress, and join discussions about Nonstationary time series modeling with applications to speech signal processing with readers worldwide — whether your club is virtual, in-person, or hybrid.

Is Readfeed free?

Yes. Creating an account and joining book clubs is free. Sign up to find readers who love the same books and start discussing today.