HG4052 · Speech Synthesis & Recognition
ə
HG4052 · Speech Synthesis & Recognition / Week 1

Inside the
Black Box.

Sounds and Numbers.

Chenzi Xu · chenzi.xu@ntu.edu.sg

Today

Roadmap

t
The shape of the morning

Today, by the clock

09:40  Lecture I: two live demos, the field, and seventy years of history
10:40  Break (10 min)
10:50  Lecture II: sounds become numbers; vectors, matrices & similarity
11:40  Break + Colab setup (10 min)
11:50  Practical: from zero Python to a vowel space you built
12:50  Wrap-up
Live demo

Two machines you already carry

DEMO 1 · ASR

Your phone takes dictation

Before we start: predict what survives. Someone dictates a sentence with a proper name and a hesitation in it. Watch what gets silently "repaired".

DEMO 2 · TTS

Your phone talks back

A screen-reader voice reads the same sentence. Listen for what it gets right effortlessly, and where it doesn't sound human.

By Week 13 you will be able to explain, and audit, both of these, level by level.

Listen closer

Two sentences that bite back

TTS · SENTENCE 1

“I live near the live music venue.”

Same four letters, two pronunciations. The voice must parse the sentence before it can say it. (How? Week 10.)

TTS · SENTENCE 2

“OK lah, see you there lor.”

The segments are easy; the meaning lives in the particles’ pitch. Listen to what the voice does to lah. (Prosody: Week 12.)

Score it as phoneticians: what does it get right effortlessly, and what is subtly wrong? Keep your notes; both sentences return in Weeks 10–12.

ʔ
Activity · predict first, then watch

Does it speak Singlish?

Before the demo: write down, word for word, what you think the recognizer will print for this sentence.

Dictated live (fallback: pre-recorded screen capture)

“Can lah, we go makan at the kopitiam after this.”

Hands up: who predicted lah survives? Who predicted makan? kopitiam?

Debrief

Whose English does it speak?

Deleted

Discourse particles (lah, lor, leh) tend to vanish: the system has no word for them.

Rewritten

makan and kopitiam get “repaired” into the nearest English the model knows.

Standardized

The output is fluent, punctuated, and someone else’s English. Fluency is not accuracy.

Our home dataset: the IMDA National Speech Corpus, thousands of hours of Singapore speech. Assignment 1 works with it; in Week 9 we audit accent bias with it. This course treats Singapore as the test set.

The wide shot

One field, six doors

PILLAR 1 · WEEKS 5–9

Speech recognition (ASR)

Speech in, text out: dictation, captions, voice search, transcription at scale.

PILLAR 2 · WEEKS 10–12

Speech synthesis (TTS)

Text in, speech out: screen readers, assistant voices, audiobooks, voice banking.

Who spoke?

Speaker recognition & diarization: voice ID, meeting minutes. W12

Where are the sounds?

Forced alignment: the corpus phonetician’s power tool. W9

Across languages

Speech-to-speech translation and voice conversion. W13

Talking back

Spoken dialogue: assistants that listen while they speak. W13

One toolbox underneath all six: features, similarity, probability, networks. That toolbox is Weeks 1–4 and 7, and it opens today.

The map

Two pipelines, mirror images

RECOGNITION (ASR): SPEECH IN, TEXT OUT Feature extraction Weeks 3–4 Acoustic model Weeks 5–8 Decoder + language model Weeks 6, 8 “the cat sat” SYNTHESIS (TTS): TEXT IN, SPEECH OUT “the cat sat” Text analysis + G2P Week 10 Acoustic model Week 11 Vocoder Week 11

This is the classical, modular map. Modern end-to-end systems (Weeks 8 & 11) fold the boxes into one network, but every job on this slide still has to happen. Evaluation, bias & ethics cut across everything: Weeks 9 & 12.

The semester

Thirteen weeks on one slide

W1
Numbers, vectors & similarity
W2
Probability & Bayes
W3
Acoustics: sampling & source–filter
W4
Features: mel & MFCCs
W5
Template matching & DTW
W6
HMMs, Viterbi & lexicons
W7
Neural networks
W8
End-to-end ASR: Whisper
W9
Evaluation, bias & alignment
W10
Classical TTS & G2P
W11
Neural TTS & vocoders
W12
Cloning, prosody & ethics
W13
Frontiers · panel · A2 poster fair
Foundations Recognition Synthesis Shaded = assessment Quiz 1 in W6 · A1 due W10 · Quiz 2 in W12 · A2 fair in W13
How you're graded

Assessment

QUIZ 1 · 20%

Week 6, in class

First 45 minutes; covers Weeks 1–5. Knowledge-based plus interpretive questions; full details nearer the quiz.

QUIZ 2 · 20%

Week 12, in class

First 45 minutes; covers Weeks 6–11. Full details nearer the quiz.

ASSIGNMENT 1 · 30%

Speech recognition

An applied ASR assignment. Released from Week 7; due Week 10. Full details in the brief.

ASSIGNMENT 2 · 30%

Speech synthesis

A synthesis assignment, presented at the Week 13 demo/poster fair. Released Week 10. Full details in the brief.

Quiz questions draw on the weekly practicals, so the notebooks are your revision. Week 13 also brings a guest panel: careers in speech technology.

How we got here · I

Two centuries of talking machines

the 2010s → 1780 1930 2010 1791 · Von Kempelen the first speaking machine 1846 · The Euphonia Faber’s talking head 1939 · The Voder Dudley’s keyboard voice W11 1951 · Pattern Playback spectrograms, played back 1953 · OVE & PAT 1961 · “Daisy Bell” a digital vocal-tract tube sings W3 1984 · DECtalk Klatt’s formant voice; later Hawking’s W10 1996 · Unit selection TTS by cut-and-splice W10 1952 · Audrey Bell Labs’ ten digits W5 1962 · IBM Shoebox sixteen words W5 1976 · Harpy (CMU) 1,011 words, DARPA W6 1980s · HMMs & HTK statistics takes over W6 1997 · Dragon NaturallySpeaking continuous dictation at home W6

Gold, above the line: machines that speak. Teal, below: machines that listen. Speaking had a 160-year head start.

The tube model returns in Week 3, the talking heads as articulatory synthesis in Week 10; the statistical turn (HMMs, and HTK, the toolkit that trained a generation) is what you build in Weeks 5 and 6. Photos: Wikimedia Commons, public domain.

How we got here · up close

1791 · The speaking machine

Von Kempelen’s own drawing (1791), via H. Traunmüller, Stockholm University; replica photo: Wikimedia Commons (PD)

  • A bellows for lungs, a reed for the voicing, a leather-and-rubber mouth shaped by one hand
  • Played like an instrument: years of practice for a few phrases, and a listener to interpret them charitably
  • No recording can exist: the phonograph was still 86 years away

Its direct descendants are the articulatory synthesizers of Week 10: in class you will play a browser vocal tract (Pink Trombone) that works on exactly von Kempelen’s plan.

How we got here · up close

1846 · The Euphonia, a talking head

Image: Wikimedia Commons (PD)

  • Joseph Faber’s machine: bellows, an ivory tongue, a rubber pharynx and a sixteen-key keyboard behind a mannequin face
  • It spoke slowly in several languages, whispered, and sang “God Save the Queen”
  • London crowds paid to hear it; a generation later, a rebuilt Kempelen machine set the young Alexander Graham Bell building his own talking head

No recording survives: the Euphonia was never captured. Machines could speak for almost a century before any machine could listen back.

How we got here · up close

1939 · The Voder goes electric

Voder demonstration, 1939 New York World’s Fair; Wikimedia Commons (PD)

  • The first electrical voice: a buzz source and a hiss source shaped by ten filter keys. Source–filter theory, built as hardware (the theory itself is Week 3)
  • Operators trained for about a year to “play” sentences by hand
  • Its sibling, Dudley’s vocoder, gave its name to the neural vocoders of W11
Hear it · “Good evening, radio audience”

Recording: D. H. Klatt’s archive, via Acoustics Today

How we got here · up close

1951 · Pattern Playback: spectrograms, played

Image: H. Traunmüller, Stockholm University

  • Paint a spectrogram, shine light through it, and hear it: analysis, run backwards
  • With it, Haskins Labs found which patterns matter: formant transitions, cue trading, categorical perception
  • Much of the speech perception you learned was discovered on this machine
Hear it · hand-painted sentences

Recording: D. H. Klatt’s archive, via Acoustics Today

How we got here · up close

1953 · OVE & PAT: the first formant synthesizers

Fant’s OVE; image: H. Traunmüller, Stockholm University

  • OVE is Orator Verbis Electris, a “speaker by electric words”; PAT, the Parametric Artificial Talker
  • No mouth at all: just a few formant tracks, steered over time
  • Their rule-driven heirs (Klatt, DECtalk) are W10, where you build vowels in a KlattGrid
OVE (Gunnar Fant, Stockholm) · “How are you? I love you!”
PAT (Walter Lawrence, UK) · “What did you say before that?”

Recordings: D. H. Klatt’s archive, via Acoustics Today

Fant is also the father of the source–filter theory you meet in Week 3: the same person gave phonetics its acoustic theory and machines their voice. The two stories are one story.

How we got here · II

2010–2020: the deep-learning decade

products & data models & methods ← 1791–2009 2010 2020 2011 · Siri ships assistants in every pocket 2014 · Amazon Echo far-field: a mic across the room 2015 · LibriSpeech the 1,000-hour open benchmark 2016 · Google Home the far-field race is on 2019 · ASR goes on-device recognition without the cloud 2011 · Kaldi HTK’s open-source heir 2012 · Deep nets take over DNNs replace GMMs in acoustic models 2014 · Deep Speech end-to-end: audio in, text out 2015 · Attention arrives Listen, Attend and Spell (seq2seq) 2017 · Human parity claimed 5.1% WER on Switchboard

The recipe behind the decade: curated data + GPUs + better algorithms. By 2020, benchmark word error rates sat below professional transcribers’.

Timeline after Hannun (2021), “The History of Speech Recognition to the Year 2030” · awni.github.io/future-speech

How we got here · III

2020–2026: the foundation-model era

machines that listen machines that speak & talk back ← the 2010s 2020 2026 2020 · wav2vec 2.0 pretraining on unlabelled speech W8 2021 · HuBERT & WavLM self-supervised representations W8 2022 · Whisper 680,000 h of web audio, 99 languages W8 2023 · MMS (Meta) ASR & TTS for 1,100+ languages W11 2024 · MERaLiON (A*STAR) Singapore’s speech LLM, raised on the NSC AUG 2026 · you are here thirteen weeks to open these boxes 2022 · AudioLM speech becomes tokens for an LM W13 2023 · VALL-E a voice cloned from 3 seconds W12 2024 · GPT-4o voice · Moshi full-duplex: it listens while it speaks W13 2025 · Voice as an interface real-time omni models: GPT-4o, Gemini, Qwen

The 2010s taught machines to hear; the 2020s fold hearing, speaking and conversing into single models. Every dot has a week on your roadmap.

New powers, new failure modes: hallucinated transcripts (Week 8), cloned voices (Week 12), accents still misheard (Week 9).

How we got here · the eras

The history in five paradigms

1950s–1980

Templates & rules

Store examples, compare distances; hand-written expert rules. Audrey, Shoebox, Harpy. W5

1980–2010

Statistics: HMM‑GMM

Speech as probability: HMMs + n‑grams. The toolkit era: HTK, then Kaldi (2011). W6

≈2010–14

Hybrid: DNN‑HMM

A deep net replaces the GMM scorer; the HMM skeleton stays. Kaldi’s neural decade. W7

≈2014–20

End-to-end

One network, audio in, text out: CTC, attention. Lexicon and HMM retired. W8

2020–now

Foundation models

Pretrain on oceans of unlabelled speech, then do everything: wav2vec 2.0, Whisper, speech LLMs. W8 · W13

End-to-end and foundation models share the architecture: one big network. What changes is the training data (oceans of unlabelled speech instead of labelled pairs for one task) and the role (one model reused for many tasks instead of one model per task).

The practicals walk the same road: Week 5 you build a template recognizer, Week 6 an HMM, Week 7 a neural classifier, Week 8 an end-to-end model. Thirteen weeks, seventy years of the field.

How we got here · a second witness

Automatic Speaker Recognition

Yu et al. 2023, Figure 1: the progress of speaker recognition systems

This is speaker recognition (“who spoke?”, W12), not ASR: and it walked the same road: spectrograms → statistical models → DNNs → pretrained models. The paradigm shifts swept every speech task at once.

Yu et al. (2023), “Twenty-Five Years of Evolution in Speech and Language Processing”, IEEE Signal Processing Magazine 40(5), Fig. 1. © IEEE, reproduced for teaching.

ɹ
Why you, specifically

The hard problems are linguistics

Coarticulation

No two /d/s are alike; every frame is context.

Variation

Accents, dialects, sociolects: systems trained on someone's speech.

Reduction

“probably” → [pɹɒbli]. Casual speech destroys dictionary forms.

Prosody

Synthetic voices read sentences, not discourse. Still unsolved.

Commercial recognizers make roughly twice the errors for Black American speakers as for white speakers (Koenecke et al. 2020; we audit this ourselves in Week 9).
The systems learned someone's phonology. Just not everyone's.

Activity · ~6 minutes, in pairs

Same sentence, five speakers

You will hear one sentence read by five speakers: Singaporean, Malaysian, Indian, British, American.

Predict, in pairs: rank the five by how many errors the recognizer will make. Commit before the reveal.

Then we compare with the actual transcripts (pre-run, so no wifi drama).

THE QUESTION UNDERNEATH

If errors track accent, the system has learned someone's phonology. The Koenecke result on the previous slide is this observation, done properly, at scale. In Week 9, you run that audit yourselves.

Break · 10 minutes · back at 10:50

Stretch. Refill.
Then: sounds become numbers.

When we return: what a recording actually is, and the first math this course runs on: vectors and similarity.

The reframe

Recorded speech is a list of numbers

A sentence (~3 s · 48,000 numbers)
Zoom: one vowel (~30 ms), the wave repeats
Zoom again: the individual samples

Sampling: measure the wave at strict, regular intervals.
16,000 of these per second is all a microphone gives the machine.
Everything else must be computed.

Live · Praat

Watch the list of numbers appear

1. Open week01_demo.wav (course repo)
2. Zoom: sentence → one vowel → single cycles
3. Keep zooming until the dots separate: samples
4. Click one dot and read its value: one number
Why this matters

The previous slide is a drawing; this is the real thing, in the tool you already know. Praat, matplotlib, and every recognizer in existence all start from these same dots.

The reframe · sampling, how often?

How many samples per cycle?

1 per cycle · the wave vanishes
2 per cycle · peak and valley only: the Nyquist limit
5 per cycle · recognisable
10 per cycle · smooth

Grey: the real wave  ·  coral dots: the samples  ·  teal: what the numbers preserve

The sampling theorem: you need more than two samples per cycle. Half the sampling rate is the Nyquist frequency, the highest frequency a recording can hold: at 16,000 samples per second, 8,000 Hz. Sound above it? Week 3: aliasing.

The reframe · quantisation, how fine?

Quantisation: every sample lands on a level

  • Between the levels nothing can be stored: each sample rounds to the nearest rung (here 8 rungs = 3 bits)
  • Playback reads the rungs back out; the gap between truth and rung is the rounding error
  • Thousands of tiny rounding errors per second add up to quantisation noise: a real, audible hiss

More bits, finer rungs, quieter hiss: 16-bit audio has 65,536 of them and the error drops below hearing. In Week 3 you listen to 16 → 8 → 4 → 3 bits and hear the rungs appear.

The reframe · quantisation, and back

Quantise in, reconstruct out

64 levels over 0–6.4 V → each level 0.1 V wide → level n starts at n × 0.1 V 64 63 62 31 30 3 2 1 0 } Reconstruction levels Original signal level = 3.09 V Playback signal level = 3.05 V Error = 0.04 V Quantisation decision levels

After Coleman (2005), Introducing Speech and Language Processing, Fig. 2.5 (adapted from Embree 1991).

  • Storage keeps only the level number: “30”, not 3.09 V. That direction is analogue to digital, written A → D: the recorder’s converter
  • Playback runs it backwards, digital to analogue (D → A): it can only output each level’s midpoint, so everything from 3.0–3.1 V comes back as 3.05 V
  • In 3.09, out 3.05: the 0.04 V gap is the quantisation error, and no later hardware can recover it

Is this the only rule? No: this figure rounds down and plays back the midpoint; the previous slide rounded to the nearest rung. Designs differ, but the guarantee is the same: the error is never more than half a level. At 16 bits the levels are 1,024 times finer than these 64, and the locked-in error sinks below hearing.

The reframe · bits

Binary: counting with 0 and 1

Place values, base 2

101₂ = 1×4 + 0×2 + 1×1 = 5

Each place is worth twice the one to its right: …8, 4, 2, 1. Add the place values that sit over a 1, and skip the rest.

EX 1.1

Use the place values to convert 1011₂, then check it in the table. How many levels do 4 bits give, and what is the largest?

Worked answers in the appendix.

All sixteen 4-bit numbers
0000 = 01000 = 8 0001 = 11001 = 9 0010 = 21010 = 10 0011 = 31011 = 11 0100 = 41100 = 12 0101 = 51101 = 13 0110 = 61110 = 14 0111 = 71111 = 15

Left half starts with 0, right half with 1: one extra bit doubles the count. That is all 2ⁿ means.

The reframe · bits, doubled

Every extra bit doubles the levels

The doubling law

1 bit → 2
3 bits → 8
8 bits → 256
16 bits → 65,536

In Week 3 you will hear what each doubling buys: roughly 6 dB less quantisation noise.

EX 1.2

8-bit audio has 256 levels. How many does 16-bit have, and how many times finer is that?

Worked answers in the appendix.

bitslevels = 2ⁿtypical rangewhere you meet it
120…1a switch: voiced / voiceless
380…7the quantisation slide’s rungs
8256−128…127telephone-era audio
1665,536−32,768…32,767speech & CD audio: int16
2416,777,216hugestudio headroom

Callbacks: the 8 rungs of the quantisation figure were the 3-bit row; real recordings live on the 16-bit row.

The reframe · data types

A data type is a reading instruction

1 · IN THE RECORDER

int16

Most recorders write 16-bit whole numbers into a WAV file. Studio kit: 24-bit; phone memos compress, then unpack.

2 · IN NUMPY

float32

Loading divides every sample by 32,768: 8192 becomes 0.25. The same wave, now decimals on one shared scale.

3 · IN THE SILICON

only bit patterns

The type says how to read a pattern: place values for integers; sign, exponent, fraction for floats.

typebitsit holdswhere you meet it
uint880…255telephone-era audio files
int1616−32,768…32,767WAV: recorders, CDs, our corpus
float3232decimals (about 7 digits)numpy audio, in [−1, 1]
float6464decimals (about 16 digits)numpy’s everyday math

After Coleman (2005), Table 2.5, translated into Python: his C types unsigned char / short / float / double are numpy’s uint8 / int16 / float32 / float64.

Why ÷ 32,768? Why [−1, 1]?

One shared, unit-free scale: 0.25 means a quarter of full loudness, whatever the file’s bit depth. And the trip is lossless: float32 holds every int16 exactly.

Overflow & headroom: why the float system wins

Mix two loud voices and you add their waves, sample by sample: 20,000 + 20,000 = 40,000, but a 16-bit box stops at 32,767. The sum does not fit: that is overflow, heard as harsh clipping. Floats have a ceiling too, but the movable decimal point (next slide) puts it near 10³⁸, and audio math never comes close: samples live in [−1, 1], so sums reach the thousands at most. That is headroom: a ceiling too far away to matter.

The reframe · inside a float

float32 is scientific notation, in binary

You have written floats since secondary school: 6.02 × 10²³ is a sign, some digits, and an exponent saying where the point sits. float32 is the same trick in base 2, budgeted into 32 bits:

1
8 exponent bits
23 fraction bits
±
where the point sits: 2ᵉ, e from −126 to 127
the digits: about 7 significant decimal digits

0.25 = +1.0 × 2⁻²  →  00111110100000000000000000000000  (the exponent field stores −2 + 127 = 125 = 01111101₂)

EXPONENT BITS BUY RANGE

8 bits give 256 positions for the point: from 2⁻¹²⁶ up to 2¹²⁷, roughly 10±³⁸. That is the headroom of the last slide.

FRACTION BITS BUY PRECISION

23 bits of digits give about 7 decimal digits. Need more? float64 re-budgets: 11 exponent + 52 fraction, about 16 digits.

Why 8 and 23? It is a budget: exponent bits buy range, fraction bits buy precision, and the split is the compromise the whole computing world standardised in 1985 (IEEE 754).

The reframe · on one slide

A recording is three choices

CHOICE 1 · SAMPLING RATE

16,000 per second

More than two samples per cycle, so 16 kHz keeps sound up to 8 kHz. Above that: gone.

CHOICE 2 · BIT DEPTH

16 bits · 65,536 levels

Every sample rounds to the nearest level; what remains of the rounding is a whisper of quantisation noise.

CHOICE 3 · DATA TYPE

int16 or float32

Whole numbers on disk; floats in numpy, rescaled to [−1, 1]. Same wave, two rulers.

The same three samples, three ways
float32  [ 0.2500,  0.1250,  0.0625 ]
int16    [  8192,   4096,   2048 ]
on disk  0010 0000 0000 0000   0001 0000 0000 0000   0000 1000 0000 0000

0.25 of full scale × 32,768 = 8192 = 2¹³: a single 1 in the 8192s place. Stored sound is these bit patterns: sixteen per sample, sixteen thousand samples per second.

The speech default: 16 kHz · 16-bit · loaded as float32. In Week 3 you make each choice audible: aliasing, bit-depth ear-training, and the industry’s standard settings.

ʃ
The course mantra

Every system in this course is a recipe for turning one list of numbers into another.

ASR: 48,000 samples → 11 characters.  TTS: 11 characters → 48,000 samples.

Math primer · I

Functions are input → output machines

spectrogram(waveform) = the picture Praat draws for you
g2p(“enough”) = /ɪˈnʌf/  (grapheme-to-phoneme: spelling → IPA)
asr(audio) = “text”

A whole recognizer is just functions composed, like rules feeding rules in a phonological derivation:

asr = decode( score( features( record(you) ) ) )

i
Math primer · II

A vowel token is a vector

Measure two formants, and a vowel becomes
a point in space (an arrow from the origin).

/i/ = (270, 2290)
      F1     F2  (Hz)

Add F3, duration, spectral tilt… same idea, more axes.
Past three dimensions you lose the picture, not the math.

vowelF1 (Hz)F2 (Hz)
i2702290
ɪ3901990
ɛ5301840
ɑ7301090
u300870

Means, adult male speakers: Peterson & Barney (1952)

You've used a vector space since first year

The vowel chart is a vector space

← F2 (Hz)  2400 ……… 700  (axis reversed, phonetician-style) ← F1 (Hz) 200 … 800 ɛæ ɔʊu i ɪ ɑ 323 Hz apart 1285 Hz apart

Phoneticians made this embedding by hand, decades before machine learning had a name for it.

Close on the chart ⇒ similar vowels.
Distance means something.

So: can we give the machine
a number for “how similar”?

+
Math primer · the moves you will reuse all term

Vector arithmetic I: moves on arrows

Adding: tip to tail
a b a + b

Walk a, then walk b from a’s tip: a + b = (2, 1) + (1, 2) = (3, 3). Add coordinate by coordinate.

Scaling: new length, same direction
a 2a

2a = (4, 2): twice as long, pointing exactly the same way. Scaling never changes direction: remember that for cosine.

Normalising: keep only the direction
v v / |v| the length-1 circle

Divide v by its own length: direction kept, size forgotten. Cosine similarity is secretly a dot product of unit vectors.

Each move is one line of numpy, and nothing here goes beyond addition and multiplication.

Math primer · the moves you will reuse all term

Vector arithmetic II: moves on vowel data

F1 (Hz) → F2 (Hz) → i ɪ (+120, −300) three /æ/ tokens their centroid: (670, 1720) (270, 2290) (390, 1990)
The difference vector: a recipe

/ɪ/ − /i/ = (390 − 270, 1990 − 2290) = (+120, −300): the instructions for getting from /i/ to /ɪ/. Careful: in raw Hz the F2 change looks bigger, but proportionally F1 moves more (+44% vs −13%). Week 4’s delta features are this idea, applied over time.

The centroid: a prototype

x̄ = (x₁ + … + xₙ) / n: average each coordinate. The centroid of your /æ/ tokens is the category’s prototype: prototype theory, as arithmetic. You compute one in the practical (and again in EX 1.5).

★ The hardest idea this week

How similar are two sounds?

Two answers, both one line of arithmetic: the dot product (angle) and Euclidean distance (gap).
Every system in this course asks this question millions of times per second.

Similarity, answer 1

The dot product asks: do they point the same way?

F1 F2 i ɪ ɑ 4.4° 27° formant vectors drawn to scale
a · b = a₁b₁ + a₂b₂

Multiply matching coordinates, add them up. What that measures: how much of one arrow lies along the other: the shadow b casts on a’s direction.

a = (a₁, 0) b = (b₁, b₂) b₂ b’s shadow on a = b₁

Lay a flat along the axis: a · b = a₁b₁ + 0 · b₂. The height b₂ never enters: only the shadow survives.

cos(a, b) = a · b / ( |a| |b| )

Cosine similarity: divide by both lengths and only direction remains: +1 = same direction, 0 = right angles. (All-positive formant vectors never get near 0.)

Bookmark this operation. In Week 3 it opens the Fourier transform: dot products with “probe” sine waves are how a computer reads the recipe of any sound.

One circle, no trigonometry homework

What cos actually measures

the other arrow's direction length 1 cos θ ≈ 0.77 θ = 40°

Cosine is the shadow a unit arrow casts on the other arrow's direction.

cos 0° = 1  same direction, full shadow
cos 45° ≈ 0.71  halfway round, shorter shadow
cos 90° = 0  right angles, no shadow at all

That is all cosine similarity uses. Multiply the shadow by both lengths and you have the dot product: a · b = |a||b| cos θ.
Week 3 preview: this circle starts spinning, and the arrow's height traces a sine wave, the blueprint of every sound.

Commit before we compute

cos( /i/, /ɪ/ ): nearer 0.9 or 0.999?

Hands up for each. Then a second guess: /ɑ/ is how many times farther from /i/ than /ɪ/ is? Write your number down.

(Most people guess the cosine too low. Watch what all-positive coordinates do to the angle.)

Worked example · by hand, once in your life

Is /i/ more like /ɪ/ or /ɑ/?

/i/ = (270, 2290)    /ɪ/ = (390, 1990)    /ɑ/ = (730, 1090)

i · ɪ = (270 × 390) + (2290 × 1990) = 105,300 + 4,557,100 = 4,662,400

lengths: |i| ≈ 2,306   |ɪ| ≈ 2,028   →   cos = 4,662,400 / (2,306 × 2,028) ≈ 0.997

i · ɑ = (270 × 730) + (2290 × 1090) = 2,693,200   →   cos ≈ 0.890

/i/ vs /ɪ/

cos 0.997 → 4.4° apart

/i/ vs /ɑ/

cos 0.890 → 27° apart

Verdict

The arithmetic agrees with your phonetics training. That's the point.

Similarity, answer 2

Euclidean distance asks: how far apart?

d(a, b) = ( (a₁−b₁)² + (a₂−b₂)² )

d(i, ɪ) = (120² + 300²)323 Hz

d(i, ɑ) = (460² + 1200²)1,285 Hz  (4× farther)

Just Pythagoras on the vowel chart: the straight-line gap between two points.

Remember this line

In Week 5 you'll compute this exact quantity between MFCC frames, thousands of times, inside your own working digit recognizer.

np.sqrt(np.sum((a-b)**2))

One line of Python. You write it in the practical, at 12:40.

Before you compute alone

The three ways this goes wrong

Crossed wires

a · b multiplies matching coordinates: a₁b₁ + a₂b₂. If you ever write a₁b₂, stop and re-pair.

The lost square root

Distance squares the gaps, then takes the root at the end. Squaring also kills minus signs: (−300)² = 90,000.

The raw-Hz trap

F2's range is roughly 3× F1's, so raw-Hz distance is mostly an F2 story. Phoneticians normalize (Bark, Lobanov) for exactly this reason.

That third one is a preview: in Week 4 the machine fixes its ruler too, bending the frequency axis to match the ear. It's called the mel scale, and you will build it yourself.

Exercises · pencils out · ~6 minutes, in pairs

Try it: vowels as vectors

EX 1.3

Peterson & Barney means: /u/ = (300, 870), /ɔ/ = (570, 840), /ɑ/ = (730, 1090). (a) Compute d(/u/, /ɔ/). (b) Which of /u/ and /ɔ/ is closer to /ɑ/? Check your verdict against the vowel chart.

EX 1.4

Let a = (1, 2) and b = (2, 4). Compute cos(a, b) and d(a, b). What does this pair of answers tell you that either alone would not?

EX 1.5

Three /æ/ tokens: (650, 1700), (670, 1740), (690, 1720). Find the centroid.

EX 1.6

One Singaporean speaker's /ɛ/, measured in Praat: (580, 1850). Peterson & Barney (US, 1952): /ɛ/ = (530, 1840), /æ/ = (660, 1720). Which is it closer to? And what can you not conclude from one token?

Worked answers in the appendix at the end of this deck.

ɑ
Why two measures?

Angle forgives, distance doesn't

A child's /i/ and an adult's /i/ sit far apart on the chart: shorter vocal tract, all formants scaled up.

But the vectors still point the same way: big distance, tiny angle.

Choosing a similarity measure is a theoretical claim about what counts as “the same sound”: speaker normalization, in your terms.

RUNNING THEME

This tension (what should the machine treat as equivalent?) returns as:

• cross-speaker failure  W5
• learned representations  W7
• accent bias in ASR  W9

M
One more object · vectors in bulk

A matrix is vectors, stacked

  • A matrix is a rectangular table of numbers; its shape is rows × columns
  • Each row here is one vowel's (F1, F2): eight vectors, one object
  • Your practical's CSV loads as exactly this table, forty rows deep
DEFINITION
Matrix

An m × n matrix is an array of numbers in m rows and n columns; the product A x applies each row's weighted sum to the vector x.

Deisenroth et al. (2020), ch. 2.
The only new operation you need this term

Matrix × vector: many dot products at once

  • Each output entry is one row's dot product with x: the operation you just learned, run once per row
The shape rule

An m × n matrix times a length-n vector gives a length-m vector. The inner sizes must agree; the outer size is what you get.

Remember this line

A @ x

One character in numpy. In Week 4 the rows become 26 mel filters; in Week 7 the same line runs with learned weights. This slide is a promise.

What a matrix does · one transform, every point

A matrix moves the whole vowel space at once

  • Multiply every vowel vector by the same matrix: the whole chart moves in one lawful sweep
  • A diagonal matrix rescales the axes: here, a child-to-adult vocal-tract scaling
  • Choosing the matrix is choosing a theory of “same vowel”: speaker normalization, as arithmetic

S has rows (1.25, 0), (0, 1.25)  ·  S · (270, 2290) = (338, 2863)

EX 1.7

Write the diagonal matrix that leaves F1 unchanged and doubles F2, then apply it to /u/ = (300, 870).

Chaining transforms

Two matrices in a row are one matrix

G has rows (2, 0), (0, 3)  ·  H has rows (1, 1), (0, 1)  ·  x = (1, 1)

first G:  G x = (2 × 1, 3 × 1) = (2, 3)

then H:  H (G x) = (2 + 3, 3) = (5, 3)

the shortcut:  H G has rows (2, 3), (0, 3), and (H G) x = (5, 3): the same answer, one multiply

H G: the same trip, one multiply G: stretch ×2, ×3 H: shear x = (1, 1) (2, 3) (5, 3)
Why it matters

A chain of matrix steps, however long, collapses into a single matrix. Linear pipelines cannot build anything one multiply could not.

Two seeds, planted

Week 4: the front end runs matrix, then log, then matrix; the log between them is exactly what keeps the two from fusing. Week 7: neural networks put a squash between layers for the same reason: without the bend, depth collapses.

Why bother, concretely

Three matrices already on your calendar

WEEK 3

The spectrogram

The picture itself is a matrix: frequencies down the rows, one 10 ms spectrum per column. You will read one before you compute one.

WEEK 4

The mel filterbank

Twenty-six rows, each a triangular ear. One multiply, filterbank @ spectrum, turns a spectrum into 26 loudnesses.

WEEK 7

A neural layer

The same multiply with weights learned from data, plus one squash: the building block of every modern speech system in this course.

EX 1.8

B has rows (1, 2), (3, 0) and (0, 1), so B is 3 × 2. Let y = (10, 5). State the shape of B y first, then compute it row by row.

EX 1.9

Week 4's filterbank is a 26 × 257 matrix and a spectrum is a length-257 vector. What is the shape of filterbank @ spectrum? And why can you not multiply them the other way round?

Worked answers in the appendix at the end of this deck.

In the textbook's words

This week, formally

DEFINITION
Vector

An ordered n-tuple of real numbers v = (v₁, …, vₙ) ∈ ℝⁿ; geometrically, a point (or an arrow from the origin) in n-dimensional space.

Deisenroth, Faisal & Ong (2020), Mathematics for Machine Learning, ch. 2.
DEFINITION
Dot product

For a, b ∈ ℝⁿ:  a · b = a₁b₁ + … + aₙbₙ = |a| |b| cos θ, where θ is the angle between a and b.

Deisenroth et al. (2020), §3.2; Jurafsky & Martin, SLP 3rd ed. (Jan 2026), §5.4.
DEFINITION
Matrix

An m × n matrix is an array of numbers in m rows and n columns; the product A x applies each row’s weighted sum to the vector x.

Deisenroth et al. (2020), ch. 2.
DEFINITION
Cosine similarity

cos(a, b) = a · b / ( |a| |b| ), ranging over [−1, 1]; for non-negative measurements (formant values, counts) the range is [0, 1].

Jurafsky & Martin, SLP 3rd ed. (Jan 2026), §5.4.
DEFINITION
Euclidean distance

d(a, b) = ( Σᵢ (aᵢ − bᵢ)² ): the straight-line distance between two points; Pythagoras, in any number of dimensions.

Deisenroth et al. (2020), §3.1.
Break · 10 minutes · back at 11:50

Open your laptop.
The Colab link is on NTULearn.

No installs, no setup: if you can open a browser tab, you are ready.

The practical · 11:50–12:50

From zero to a vowel space you built

  • Run your first Python, in the browser, nothing to install
  • Loops & lists over real formant data
  • numpy: means, difference vectors
  • Plot the IPA vowel chart from numbers
  • Cosine similarity & distance: today's lecture, executable

Stretch (or take-home): a nearest-neighbour “vowel guesser”: drop in your own Praat measurements and see where your accent lands.

You leave with this, made by you
iɪɛ æɑɔ ʊu matplotlib, axes flipped phonetician-style
The practical · your first five minutes, scripted

Nothing to install, one thing to break

1. Open the Week 1 Colab link (NTULearn → Practicals)
2. Run the first cell: Shift + Enter
3. Run the data cell: one wget fetches every file we use today
4. Run the cell that is designed to fail, and read the error aloud with me

Step 4 fails on purpose. An error message is the computer telling you what went wrong and where, not a judgement on you. You will meet many this term; we read the first one together so they never feel scary.

The practical · survival kit I

Paths: every file has an address

Read it right to left

/content/data/vowels.csv

A file vowels.csv, inside a folder data, inside a folder content, inside / : the root, the outermost folder of the whole machine.

ABSOLUTE PATH

/content/data/vowels.csv
Starts with / : directions from the root. Works no matter where you stand.

RELATIVE PATH

data/vowels.csv
No leading slash: directions from where you are standing. .. is the folder above; . is here; ~ is your home folder.

Which slash?

Mac and Linux write paths with the forward slash /. Windows uses the backslash \ on your own laptop. Here it will never matter: a Colab notebook runs on a Linux machine in the cloud, so whatever laptop you bring, everyone types /.

The practical · survival kit II

Three commands: pwd, ls, cd

$ pwd
/content            ← where am I standing?
$ ls                  ← what is here?
data  sample_data
$ cd data             ← step into data
$ pwd
/content/data
$ cd ..                ← up one level
$ cd -                 ← back where I just was
$ cd /                 ← jump to the root
$ cd ~                 ← jump to your home folder
In Colab

Prefix shell commands with ! : !pwd, !ls (you already ran !wget today). One catch: !cd does not stick: every ! line opens a fresh shell. To actually move, use %cd data.

Why learn this

In Week 9 you drive a real command-line tool, the Montreal Forced Aligner, with exactly these three commands.

Your files in Colab

Downloads land in /content, which is wiped when the runtime shuts down. To keep work, mount your Drive: drive.mount('/content/drive') → your files appear under /content/drive/MyDrive/.

On your own machine (optional)

Mac: the Terminal app (Applications → Utilities, or Spotlight “Terminal”). Windows: PowerShell or Windows Terminal: pwd, ls and cd work there too (the older cmd.exe spells things differently). Nothing in this course requires a local terminal: Colab gives everyone the same Linux shell in the browser.

The practical · the hour, mapped

By the end of the hour, you will have built this

1.  Orientation: paths, pwd / ls / cd, the deliberate error (15 min)
2.  Lists & loops over real formant data (10 min)
3.  numpy: means & difference vectors (10 min)
4.  Plot the vowel space, axes flipped (15 min)
5.  Similarity two ways: cosine & distance (10 min)
DONE LOOKS LIKE

✓ a vowel-space plot, IPA labels, axes reversed
✓ every vowel ranked by similarity to /i/
✓ one error message you fixed yourself
✓ pwd, ls and cd in your fingers

The vowel-guesser stretch is take-home if the hour runs out. This notebook's skills return in Assignment 1.

The course book

One free book underneath it all

Speech and Language Processing

Daniel Jurafsky & James H. Martin
3rd edition (online draft)

web.stanford.edu/~jurafsky/slp3

Free chapter PDFs. Nothing to buy, ever.

  • The field’s standard textbook, and it is free online
  • A living draft, updated continually: we cite the January 2026 release, and chapter numbers move between releases, so follow the syllabus’s numbers
  • Our chapters: Ch 3 n‑grams (W2, W6) · Ch 14 phonetics & features (W3–4) · Ch 15 ASR (W5–9) · Ch 16 TTS (W10–12) · Ch 25 conversation (W13)
Before next week

This week’s readings

REQJurafsky & Martin, Speech and Language Processing (3rd ed., free online): Ch 15 & 16, intros only

Just the chapter openings of the ASR and TTS chapters: the bird's-eye view of today's two pipelines. web.stanford.edu/~jurafsky/slp3 (chapter numbers refer to the Jan 2026 release)

REQ3Blue1Brown, Essence of Linear Algebra: “Vectors”, “Dot products”, “Linear transformations and matrices” & “Matrix multiplication as composition” (~45 min video)

The geometric picture from today, animated better than any slide; the last two episodes are today's matrix slides, moving. 3blue1brown.com/lessons/ vectors · dot-products · linear-transformations · matrix-multiplication

OPTVanderPlas, A Whirlwind Tour of Python (free)

Keep open as a reference during practicals; variables, lists, loops are all we use for two weeks.

OPTDeisenroth, Faisal & Ong, Mathematics for Machine Learning (free PDF) · opening of Ch 2

Only for the math-hungry: the formal version of today's definitions. mml-book.github.io

Looking one week ahead

Next week’s readings

REQGoldsmith (2007), “Probability for linguists”

Written exactly for this audience: probability taught through phones and words, no engineering background assumed. Free online.

REQJurafsky & Martin, SLP Ch 3 “N-gram Language Models”: §3.1–3.3 and the perplexity section

Skip smoothing for now: it returns in Week 6. (Chapter numbers refer to the Jan 2026 release.)

REQ3Blue1Brown, “Bayes theorem, the geometry of changing beliefs” (15 min video)

Next week’s hardest idea, animated: watch it before Friday and the lecture will feel like revision.

OPTShannon (1951), “Prediction and Entropy of Printed English”

The paper behind Week 2’s guessing game; surprisingly readable.

OPTSeeing Theory (Brown University)

Interactive probability, almost no prose: play with it for ten minutes.

Next week: probability & Bayes’ rule, the math of a listener’s guesses. “wreck a nice beach” awaits.

Wrap-up · quiz radar

Six things to walk out with

Exit ticket before you leave: one sentence, what is still muddiest? (form link on NTULearn)

ə

One list of numbers
into another.

That's the whole course. See you in the lab.

Sources

References

Coleman, J. (2005). Introducing Speech and Language Processing. Cambridge University Press. Ch. 2 “Sounds and numbers”: sampling, quantisation & reconstruction (Fig. 2.5), data types (Table 2.5).

Hannun, A. (2021). “The History of Speech Recognition to the Year 2030.” awni.github.io/future-speech. Basis of the 2010–2020 timeline.

Traunmüller, H. “Wolfgang von Kempelen’s speaking machine and its successors.” Stockholm University. resources.ling.su.se/hartmut/kemplne.html. The early-machines tour and images.

Klatt, D. H. (1987). “Review of text-to-speech conversion for English.” Journal of the Acoustical Society of America 82(3), 737–793. Historical recordings from Klatt’s archive, via Acoustics Today (acousticstoday.org).

Yu, D. et al. (2023). “Twenty-Five Years of Evolution in Speech and Language Processing.” IEEE Signal Processing Magazine 40(5), 27–39. Fig. 1 reproduced.

Koenecke, A. et al. (2020). “Racial disparities in automated speech recognition.” PNAS 117(14), 7684–7689.

Peterson, G. E. & Barney, H. L. (1952). “Control methods used in a study of the vowels.” JASA 24(2), 175–184. The formant means used throughout this deck.

Images: Wikimedia Commons (public domain) and H. Traunmüller.  Recordings: D. H. Klatt’s history-of-synthesis archive, via Acoustics Today.

Appendix · worked answers

Answers · I

EX 1.1

1011₂ = 8 + 0 + 2 + 1 = 11 (check: row 1011 of the table). Four bits give 2⁴ = 16 levels; the largest is 1111₂ = 15.

EX 1.2

2¹⁶ = 65,536 levels: 65,536 ⁄ 256 = 256 times finer. Each extra bit doubles the levels and buys about 6 dB less noise (Week 3).

EX 1.3

(a) Δ = (270, 30) → d = (72 900 + 900) = 73 800272 Hz.
(b) d(/u/, /ɑ/) = (430² + 220²) ≈ 483 Hz;  d(/ɔ/, /ɑ/) = (160² + 250²) ≈ 297 Hz → /ɔ/ is closer, exactly what the chart shows.

EX 1.4

a · b = 2 + 8 = 10; |a| = 5, |b| = 20, |a||b| = 10 → cos = 1 (perfectly collinear). But d = (1² + 2²) = 52.24.
Same direction, different magnitude: the angle says “identical quality”, the distance says “different point”. The small-vs-large vocal tract situation, in miniature.

EX 1.5

Centroid = ( (650+670+690)/3 , (1700+1740+1720)/3 ) = ( 2010/3 , 5160/3 ) = (670, 1720): the mean vector, component by component. (Exactly how we summarise a speaker's vowel space, and how k-means will later find clusters.)

Appendix · worked answers

Answers · II

EX 1.6

d(to /ɛ/) = (50² + 10²)51 Hz; d(to /æ/) = (80² + 130²)153 Hz → closer to P&B /ɛ/. What you cannot conclude: anything about Singapore English generally. One token from one speaker is an anecdote; in Week 3 you measure your own vowels, and in Week 9 you learn to do this at corpus scale.

EX 1.7

Rows (1, 0) and (0, 2). Applied to /u/: 1×300 + 0×870 = 300; 0×300 + 2×870 = 1740, so the result is (300, 1740): F1 untouched, F2 doubled. Diagonal entries are per-axis volume knobs; the off-diagonal zeros mean the axes do not mix.

EX 1.8

Shape first: B is 3 × 2 and y has length 2, so B y has length 3. Rows: 1×10 + 2×5 = 20;  3×10 + 0×5 = 30;  0×10 + 1×5 = 5. So B y = (20, 30, 5): three weighted sums, one object. In Week 4 the rows are mel filters and the three becomes twenty-six.

EX 1.9

(26 × 257) times a length-257 vector: inner sizes agree, so the result has length 26: one loudness per mel filter. The other way round fails the shape rule: a length-257 vector against 26 rows has no matching inner size. Order matters in matrix multiplication, and the shapes tell you which order is meaningful.