All videos
0:00 / 0:00
research

Formant Synthesis, Concatenative Synthesis & Statistical Methods for TTS

Valerio Velardo - The Sound of AI16 June 2026Watch on YouTube

Part of series

Ep. 2 · Cloning Speech Text

View the series

Description

Learn about traditional text-to-speech techniques before the rise of neural networks in 2016. Explore formant synthesis, concatenative synthesis, and statistical parametric (HMM-based) synthesis—the methods that paved the way for modern neural TTS. This is video 5 in The Monster Text-to-Speech and Voice Cloning Course, a lecture series designed to give you a deep understanding of state-of-the-art concepts in speech synthesis. 🎯 KEY TOPICS: - The evolution of speech synthesis before deep learning - How formant synthesis modeled the vocal tract - How concatenative synthesis stitched recorded speech units - The rise of HMM-based (parametric) synthesis - Why pre-neural voices sounded robotic or over-smoothed - How these classic methods paved the way for neural TTS CONSULTING: 🚀 AI Music + Audio Consulting: https://valeriovelardoadvisor.com/ 📩 Get my AI Music content in your inbox for free: https://valeriovelardo.substack.com/ COURSE MATERIALS + DISCUSSION: - GitHub Repository: https://github.com/musikalkemist/tts-voicecloning-course - Join The Sound of AI Slack Community: https://valeriovelardo.com/the-sound-of-ai-community/ (#tts-course channel) Content: 0:00 Intro 4:05 Formant synthesis 7:41 Formant: Pros and cons 12:18 Concatenative synthesis 13:41 Diphone concatenation 15:10 Unit selection 25:20 Concat: Pros and cons 27:55 Statistical parametric synthesis (HMM) 38:57 HMM-based TTS: Pros and cons 42:32 Comparing traditional TTS

What you'll learn

  • Formant synthesis models the vocal tract using formants to generate speech, but produces unnatural sounds because it lacks suprasegmental features and natural expression.
  • Concatenative synthesis stitches together recorded speech fragments (diphones and units), yielding more natural results but creating artifacts at the boundaries between units.
  • HMM-based (parametric) synthesis combines statistical models with acoustic parameters and offers greater control, yet results in over-smoothed, unnatural-sounding speech.
  • These classical methods laid the foundation for neural TTS by revealing their limitations and demonstrating the need for improved approaches.

Frequently asked questions

How does formant synthesis work and what are its main drawbacks?
Formant synthesis models the vocal tract with formants (resonances) to artificially generate speech. Its main drawbacks are unnatural-sounding output and lack of suprasegmental features like prosody and expression.
What is the difference between concatenative synthesis and HMM-based synthesis?
Concatenative synthesis stitches together recorded speech fragments for more natural sound, while HMM-based synthesis uses statistical models for greater control. Concatenation creates artifacts at seams, HMM synthesis produces over-smoothed output.
Why were classical TTS methods replaced by neural networks?
The classical methods had clear limitations: formant synthesis sounded unnatural, concatenation produced artifacts, and HMM synthesis was over-smoothed. Neural networks offered a better way to generate more natural-sounding speech.
What is diphone concatenation and how is it used in concatenative synthesis?
Diphone concatenation focuses on joining speech fragments between two phonemes (diphones). This helps make transitions between sounds more natural, though artifacts can occur where units are spliced together.

Topics