
This course introduces the linguistic, acoustic, and computational foundations of automatic speech recognition (ASR) and text-to-speech (TTS) synthesis. The course is designed for linguistics students: technical content is mostly taught conceptually and visually, with the mathematical foundations built in Weeks 1–2 and a deliberate emphasis on phonetic intuition over mathematical derivation. Students move from the speech chain and acoustic phonetics, through digital signal representations and classical statistical systems (dynamic time warping, HMMs, n-grams, unit selection), to modern neural architectures (CTC, attention, Whisper, neural vocoders, voice cloning), treated as systems whose inputs, outputs, and failure modes can be analysed phonetically. Particular attention is paid to Singapore’s multilingual speech environment: Singlish, accent variation, Mandarin–English code-switching, and the IMDA National Speech Corpus. Hands-on work uses Praat and scaffolded Google Colab notebooks (librosa, Whisper, neural TTS models). Ethical issues are engaged where they arise — recognition bias, voice cloning and consent, low-resource languages — and the course closes with a guest panel on careers in speech technology and a poster fair.
Fridays 09:30 – 13:20, TR+69 · Prerequisite: HG2003 Phonetics & Phonology. No programming experience required — all Python work runs in scaffolded Google Colab notebooks.