Speech Synthesis and Recognition

Abstract

This course introduces the linguistic, acoustic, and computational foundations of automatic speech recognition (ASR) and text-to-speech (TTS) synthesis. The course is designed for linguistics students: technical content is mostly taught conceptually and visually, with the mathematical foundations built in Weeks 1–2 and a deliberate emphasis on phonetic intuition over mathematical derivation. Students move from the speech chain and acoustic phonetics, through digital signal representations and classical statistical systems (dynamic time warping, HMMs, n-grams, unit selection), to modern neural architectures (CTC, attention, Whisper, neural vocoders, voice cloning), treated as systems whose inputs, outputs, and failure modes can be analysed phonetically. Particular attention is paid to Singapore’s multilingual speech environment: Singlish, accent variation, Mandarin–English code-switching, and the IMDA National Speech Corpus. Hands-on work uses Praat and scaffolded Google Colab notebooks (librosa, Whisper, neural TTS models). Ethical issues are engaged where they arise — recognition bias, voice cloning and consent, low-resource languages — and the course closes with a guest panel on careers in speech technology and a poster fair.

Date
Aug 10, 2026 — Nov 13, 2026
Event
HG4052 · Semester 1, AY 2026/27
Location
TR+69, Nanyang Technological University, Singapore

Fridays 09:30 – 13:20, TR+69 · Prerequisite: HG2003 Phonetics & Phonology. No programming experience required — all Python work runs in scaffolded Google Colab notebooks.

Weekly slides

Week 1 10 – 14 Aug Inside the black box: what ASR and TTS do · Math primer I: vectors & similarity View slides → Week 2 17 – 21 Aug Math primer II: probability, Bayes’ rule, logs & surprisal View slides →
Week 3 24 – 28 Aug Acoustics primer: sampling, aliasing, the Fourier idea, spectrograms, source–filter
Week 4 31 Aug – 4 Sep From waveform to features: mel spectrograms, MFCCs, F0
Week 5 7 – 11 Sep ASR I: variability, template matching, dynamic time warping, word error rate
Week 6 14 – 18 Sep ASR II: HMMs, Viterbi, lexicons & n-gram language models · Singlish case study
Week 7 21 – 25 Sep Neural networks primer: from neurons to vowel classifiers
Recess week · 28 Sep – 4 Oct
Week 8 5 – 9 Oct End-to-end neural ASR: CTC, attention, wav2vec 2.0, Whisper
Week 9 12 – 16 Oct Evaluation, bias & forced alignment · speech corpora and ethics
Week 10 19 – 23 Oct TTS I (classical): text normalisation, G2P, formant / unit-selection / articulatory synthesis
Week 11 26 – 30 Oct TTS II (neural): acoustic models, vocoders, end-to-end voices, speaker embeddings
Week 12 2 – 6 Nov Prosody, voice cloning & the ethics of synthetic voices
Week 13 9 – 13 Nov Frontiers: spoken dialogue, speech LLMs, speech-to-speech translation · careers panel & poster fair
Dr Chenzi Xu
Dr Chenzi Xu
Assistant Professor

My research interests include speech prosody, speech perception, and speech technology.

Previous

Related