Abstract
This course introduces the linguistic, acoustic, and computational foundations of automatic speech recognition (ASR) and text-to-speech (TTS) synthesis. The course is designed for linguistics students: technical content is mostly taught conceptually and visually, with the mathematical foundations built in Weeks 1–2 and a deliberate emphasis on phonetic intuition over mathematical derivation. Students move from the speech chain and acoustic phonetics, through digital signal representations and classical statistical systems (dynamic time warping, HMMs, n-grams, unit selection), to modern neural architectures (CTC, attention, Whisper, neural vocoders, voice cloning), treated as systems whose inputs, outputs, and failure modes can be analysed phonetically. Particular attention is paid to Singapore’s multilingual speech environment: Singlish, accent variation, Mandarin–English code-switching, and the IMDA National Speech Corpus. Hands-on work uses Praat and scaffolded Google Colab notebooks (librosa, Whisper, neural TTS models). Ethical issues are engaged where they arise — recognition bias, voice cloning and consent, low-resource languages — and the course closes with a guest panel on careers in speech technology and a poster fair.
Date
Aug 10, 2026 — Nov 13, 2026
Event
HG4052 · Semester 1, AY 2026/27
Location
TR+69, Nanyang Technological University, Singapore
Fridays 09:30 – 13:20, TR+69 · Prerequisite: HG2003 Phonetics & Phonology. No programming experience required — all Python work runs in scaffolded Google Colab notebooks.
Weekly slides
Week 1
10 – 14 Aug
Inside the black box: what ASR and TTS do · Math primer I: vectors & similarity
View slides →
Week 2
17 – 21 Aug
Math primer II: probability, Bayes’ rule, logs & surprisal
View slides →
Week 3
24 – 28 Aug
Acoustics primer: sampling, aliasing, the Fourier idea, spectrograms, source–filter
View slides →
Week 4
31 Aug – 4 Sep
From waveform to features: mel spectrograms, MFCCs, F0
View slides →
Week 5
7 – 11 Sep
ASR I: variability, template matching, dynamic time warping, word error rate
View slides →Week 6
14 – 18 Sep
ASR II: HMMs, Viterbi, lexicons & n-gram language models · Singlish case study
Week 7
21 – 25 Sep
Neural networks primer: from neurons to vowel classifiers
Recess week · 28 Sep – 4 Oct
Week 8
5 – 9 Oct
End-to-end neural ASR: CTC, attention, wav2vec 2.0, Whisper
Week 9
12 – 16 Oct
Evaluation, bias & forced alignment · speech corpora and ethics
Week 10
19 – 23 Oct
TTS I (classical): text normalisation, G2P, formant / unit-selection / articulatory synthesis
Week 11
26 – 30 Oct
TTS II (neural): acoustic models, vocoders, end-to-end voices, speaker embeddings
Week 12
2 – 6 Nov
Prosody, voice cloning & the ethics of synthetic voices
Week 13
9 – 13 Nov
Frontiers: spoken dialogue, speech LLMs, speech-to-speech translation · careers panel & poster fair

Assistant Professor
My research interests include speech prosody, speech perception, and speech technology.