HG4052 · Speech Synthesis & Recognition
Σ
HG4052 · Speech Synthesis & Recognition / Week 8

The End of the
Assembly Line.

Blanks and Attention.

Chenzi Xu · chenzi.xu@ntu.edu.sg

t
The shape of the morning · welcome back from recess

Today, by the clock

09:40  Welcome back: five echoes, then Assignment 1 (15 min)
09:55  Lecture I: limits of the classical pipeline, and CTC
10:40  Break (10 min)
10:50  Lecture II: attention, then Whisper on Singlish (model details: extra)
11:40  Break + Colab setup (10 min)
11:50  Practical: the CTC collapse in code, then Whisper on Singlish
12:50  Wrap-up
Across the recess · 3 minutes, tell your neighbour

Five echoes from Weeks 6 and 7

1. The fundamental equation of classical ASR?

Ŵ = argmaxW P(X | W) · P(W): the acoustic likelihood times the language-model prior

2. DTW and Viterbi fill the same kind of table. What differs in the cells?

DTW keeps the minimum cost of the cells leading in; Viterbi keeps the maximum probability

3. What does a softmax output layer produce?

one probability per class: positive numbers that sum to one

4. How does a network learn its weights?

the loss (cross-entropy) scores its mistakes; gradient descent moves each weight a step downhill

5. In the hybrid recognizer, what did the network replace, and what stayed?

it replaced the Gaussian emission scores; the HMM, lexicon, Viterbi and n-gram language model stayed

Today one network replaces every part of echo 1.

A
Assignment 1 · two weeks to go

Assignment 1: where you should be

  • Brief read; data-access request sent. Not sent yet? Send it today: approval takes days
  • Today's practical runs steps 2 to 4 of the brief (system A, the normalizer and WER, system B) on two clips; you repeat them on your own test set
  • Due Friday 23 October, the end of Week 10
KEY DATES

Released: 25 Sep (Week 7)
Due: Fri 23 Oct (Week 10)
Weight: 30%

Everything task-specific is in the brief on NTULearn.

Δ
Today

Roadmap

In the practical: write the CTC collapse, then test Whisper on Singlish speech.

The classical pipeline, Weeks 4 to 7

Three limits of the classical pipeline

MFCC framesfeatures
→
lexiconwords → phones
→
HMM acoustic modelViterbi
→
n-gram priorword sequences
→
best sentence
LIMIT 1

Modules trained separately

  • Each module has its own objective (acoustic likelihood, n-gram counts, a hand-written lexicon)
  • None is trained on transcript errors: someone must find which module caused an error and fix it by hand (add a lexicon entry, adapt the acoustic model)
LIMIT 2

A fixed lexicon

  • A word missing from the lexicon (“makan”) cannot be output
  • Someone must add it by hand
  • A third hand-built model, beside the acoustic model and the language model
LIMIT 3

Frames scored one at a time

  • Given its state, each frame is scored on its own
  • Coarticulation (Week 3) ties neighbouring frames together
  • The hybrid's window of frames eases this but does not remove it

End-to-end recognition: train one network on (audio, text) pairs; it learns the acoustic model, the spelling and the word sequences together.

Limit 3, drawn · illustrative scores

The hybrid recognizer scores each frame on its own

  • The same network runs once per frame: for x6 it reads a short window (x4 to x8) and gives a probability for each state
  • Each column is scored on its own: x6's scores do not use the scores given to x5
  • Only the HMM's transitions (stay, or move on) link neighbouring frames, inside Viterbi
The alignment problem

Two seconds of audio. Two letters of text.

  • A 2-second clip of “no” is 200 frames (one every 10 ms); the transcript is 2 characters
  • Which frames belong to which letter? This is the alignment problem
  • Weeks 5 and 6 solved it with a hard path through a table: DTW, then Viterbi
  • A single network outputs scores frame by frame; nothing inside it picks a path. The field found two answers
ANSWER 1 · CTC

Do not choose one alignment: add up the probabilities of all of them.

ANSWER 2 · ATTENTION

At each output step, compute a weight for every frame: a soft alignment, in which every frame counts a little.

★ The hardest idea this week

Two hundred frames. Two letters.
Don't choose. Sum.

Connectionist temporal classification (CTC): the network labels every frame,
a collapse rule turns the labels into a transcript, and training adds up the probabilities of every labelling that collapses to the correct one.

Two rules, one new symbol

The blank, and the collapse

  • Every frame gets a label: a character, or the blank ∅, meaning “no new letter at this frame”
  • Collapse rule, applied in order: merge repeats, then delete blanks
  • The blank has two jobs: it labels frames that start no new letter (silence, or the middle of a sound), and it keeps real double letters apart

Spot the difference: b o o k collapses to bok; b o ∅ o k collapses to book. Without a blank between them, repeats always merge. The ∅ needs no silence: in b b o o ∅ o k it sits inside the vowel.

Worked collapses

n n ∅ o o  → n ∅ o  → no
∅ n o o ∅  → ∅ n o ∅  → no
n ∅ n o  → n ∅ n o  → nno

many frame strings spell “no”; one of these does not

CTC, read as a phonetician

Frame labels as a narrow transcription

  • Illustrative frame labels, one per 10 ms: k k ∅ æ æ æ ∅ t t
  • Merge repeats: k ∅ æ ∅ t; delete blanks: kæt
  • The collapse moves from narrow to broad: timing and length go, the sequence of categories stays
  • Training rewards every frame string that collapses to the target, just as [kʰæːt] and [kʰætʰ] both count as /kæt/
Three frame strings, one broad form

k k ∅ æ æ æ ∅ t t  → kæt
k ∅ æ æ ∅ ∅ t  → kæt
∅ k æ ∅ t t ∅  → kæt

phone labels here; English CTC models usually output letters

Decoding · illustrative numbers

Greedy decoding: eight frames to “no”

  • A network ends in a softmax over {n, o, ∅} at every frame (a toy alphabet; English CTC models use about 30 symbols): the bars show those probabilities, three per frame
  • Greedy decoding: per frame, take the most probable label
  • Then the collapse: merge repeats, delete blanks. Eight frame labels become two letters, with no lexicon and no Viterbi

Greedy decoding can miss the most probable transcript; beam search (Week 6) often finds it. See the take-home stretch.

Training: the table trick, again · illustrative numbers

Don't pick a path. Add them all.

  • Many frame strings collapse to “no”; CTC defines P(“no” | audio) as the sum over all of them
  • A path is one label per frame: 38 = 6,561 paths. Its probability is its 8 frame probabilities multiplied (the greedy path: 0.047)
  • 210 of these paths spell “no”; together P = 0.280
  • The loss is −log of that sum (Week 7's cross-entropy)
  • One table, filled once, adds them all: the forward table
Every possible transcript · training needs only “no”
transcriptpathstotal P
………
………
………
………

…and 105 more: all 109 transcripts add up to 1; “no” wins on probability, not on number of paths.

THE TABLE TRICK, AGAIN

Week 5, DTW: the minimum over paths.
Week 6, Viterbi: the maximum; the forward algorithm: the sum.
Week 8, CTC: the forward sum; its −log is the training loss.

J&M §16.5.2.

Training: computing the sum · the same eight frames

Setting up the forward table

Training: computing the sum · the same eight frames

Filling the table, one cell at a time

new cell = ( same row + one row up + two rows up* )previous frame × P(row's label)this frame

* only to skip a ∅ between two different letters. Never three rows up: that would jump over a letter.

  • Outlined cells at x8: all 210 paths, 0.006 + 0.274 = 0.280; loss = −log 0.280 = 1.27
  • Training changes the network's weights to raise this total (Week 7)
Limits of CTC

What CTC still assumes

Labels independent, given the audio

Each frame's label is chosen without looking at the labels on other frames. J&M §16.5.1: “CTC makes a strong conditional independence assumption”

No language model inside

CTC spells by sound: a real CTC model wrote “makan” as MACKIN. Systems add a separate language model at decoding time, or use an attention decoder

Left to right only

Labels follow the audio in order: suitable for speech, not for translation, where word order changes

At least one frame per letter

U letters need at least U frames, plus a ∅ between double letters: “book” needs 5 (b o ∅ o k)

k
10:40 · Break

CTC: every alignment, summed.
Next: alignment, learned.

After the break: attention, then wav2vec 2.0 and Whisper.

The second answer

Encoder and decoder

α
Attention, step by step

Attention is a weighted sum, learned

score each encoder vector against the decoder's state (a dot product)  →  softmax into weights  →  weighted sum

  • The weights: a probability distribution over frames, saying which frames matter for this token
  • The weighted sum: a matrix × vector product, as in Week 4's filterbank
  • New weights for every token; training learns the scoring
SELF-ATTENTION
  • The same step inside the encoder: each frame scores every frame
  • Each frame's new vector: a weighted sum over all frames, so it carries its context (coarticulation, limit 3)

Transformer: a stack of blocks; each block is self-attention, then feedforward layers. The 3Blue1Brown attention video shows one step by step.

Illustrative weights, not from a trained model · four letters, one clip of “kopi”

Attention weights for “kopi”

  • Each row: one decoder step's weights over the frames; each row sums to 1
  • In trained models the high weights move left to right; no rule forces this order: it comes from the speech data
  • Compare DTW (Week 5): a similar diagonal, fixed there by a rule, learned here by gradient descent on transcripts
  • An attention plot looks like a forced alignment, Week 9's tool
Simplified illustration · where the weights come from

Attention is learned from the transcript alone

  • Training data: the audio and the transcript “kopi”; nobody marks which frames belong to which letter
  • Epoch 0: every frame gets the same weight, so each letter's prediction is a blur and the loss is high
  • Gradient descent (Week 7) raises the weights on frames that help predict the correct letter: each row sharpens onto its own sound

Simplified: here each letter has its own row of scores for this one clip. A real network learns the scoring itself (the matrices behind the dot products) from thousands of clips, so it can align clips it has never seen.

+
Pretrained models · what is core, what is extra

wav2vec 2.0 and Whisper in two lines

WAV2VEC 2.0

Learns from unlabelled audio first, then from transcribed speech (1 hour can be enough), with CTC

WHISPER

An encoder-decoder with attention, trained on 680,000 hours of web audio: the model in today's practical

Extra for the curious · learning from unlabelled speech

wav2vec 2.0: pretraining on unlabelled audio

  • Labels are the bottleneck: transcribing speech is slow and costly, and most transcripts exist for a few large languages
  • So pretrain without labels: hide some frames; the network must pick each hidden frame's true sound code out of decoys from the same clip
  • Compare infants, who form phonetic categories from the sounds they hear before they know words (Maye, Werker & Gerken 2002)
  • Then fine-tune (continue training on labelled pairs) with a CTC output layer
WHY IT MATTERED

Pretrained on 53,000 unlabelled hours, then fine-tuned: 1 hour of labels beat the previous best, which had 100 hours of labels; 10 minutes of labels still gave 4.8% WER on clean read English (with a language model). Good news for languages with few transcripts.

Baevski et al. (2020), NeurIPS 33, 12449–12460. J&M §16.4 describes HuBERT, a close relative.

Extra for the curious · simplified drawing

wav2vec 2.0: pretrain, then fine-tune

To learn more: Boigne (2021), “An Illustrated Tour of Wav2vec 2.0”. jonathanbgn.com/2021/09/30/illustrated-wav2vec-2.html

W
Extra for the curious · weak supervision, at scale

Whisper: architecture and training data

?
Commit before you see it · hands up

Whisper hears “can lah, we go makan”. What does it print?

A  the transcript, word for word, with the particles and “makan”

B  something fluent but anglicised: “Can, let's go to McCann”

C  nothing: the audio is not in its training data

Whisper's results

What 680,000 hours gave, and what they did not

Gained
  • Robust to noise
  • Far fewer errors than models trained on LibriSpeech (English audiobooks), on 12 other test sets, without fine-tuning
  • Punctuation and capital letters included
  • No lexicon to maintain
Still weak
  • Singlish particles and code-switching: not tested in the paper
  • On silence or long pauses: fluent text nobody said (hallucination)
  • No lexicon file to fix “makan” by hand: the remaining option is better training data
A COMMUNITY FINE-TUNE
  • whisper-small fine-tuned on ~122,000 NSC Part 2 samples: prompted read speech, ~161 speakers
  • Model card: 9.69% WER on held-out Part 2 prompts (~43,000 samples); no base-model score
  • Training and test come from the same corpus, so the number flatters it
  • Out of scope, by its own card: conversation

jensenlwt, whisper-small-singlish-122k (Hugging Face model card); Radford et al. (2023), §6; Koenecke et al. (2024).

What changed for linguists

What linguists lost and gained

CLASSICAL PIPELINE · LOST

Every part can be inspected · a bad pronunciation is fixed by editing one lexicon line · phonetic knowledge entered each module directly

END TO END · GAINED

Lower error rates, given enough training data · coarticulation modelled, not assumed away · new languages without hand-built lexicons · one model to deploy

The linguist's role now:

V
11:40 · Break, then the practical

Run Whisper.
Then find its errors.

Open the Week 8 notebook (link on NTULearn). In the next hour: the CTC collapse, whisper-small on three clips, and a base versus fine-tuned comparison on one Singapore English clip.

The practical

The CTC collapse, then Whisper on Singlish

  • Code first: finish the CTC collapse (one line) and decode the lecture's eight-frame grid
  • Two models, whisper-small (244 million parameters) and its Singlish fine-tune, on two clips: clean read speech, and your own voice saying a Singapore English sentence (a synthetic fallback is provided)
  • The A/B: lowercase and strip punctuation on both sides, then compute WER for each model
  • Read the errors phonetically: the particle, the borrowed words, the vowels: which survive fine-tuning?
  • Silence test: input five seconds of silence; record what each model outputs; write a two-line diagnosis

Stretch (take-home): a two-frame grid with P(a) = 0.4 and P(∅) = 0.6 on both frames. Greedy reads ∅ ∅ and outputs nothing, yet the three paths that spell “a” add up to 0.64. Explain why a search over transcripts picks “a”. Illustrative numbers.

You leave with this, made by you

illustrative layout: one clip, two models, two transcripts, two WERs: your first audit table

The practical hour

The practical hour, stage by stage

1.  The CTC collapse TODO; decode the lecture's grid (15 min)
2.  Two models on two clips: clean read speech and your own voice (20 min)
3.  The A/B: normalize, then WER for both models (15 min)
4.  Silence test: what each model outputs; two-line diagnosis (5 min)

The stretch is take-home; stages 1 to 4 are the in-class core.

YOU LEAVE WITH

A collapse function you finished, two models' transcripts of two clips, a WER table, and a record of what Whisper does with silence.

If a model download stalls, the notebook uses stored transcripts of the provided clips.

In the textbook's words

This week, formally

DEFINITION
Connectionist temporal classification

The intuition of CTC is “to output a single character for every frame of the input”; a function then “collapses all repeated letters and then removes all blanks”. Course gloss: P(transcript | audio) is the sum over every frame-label path that collapses to it.

Jurafsky & Martin, SLP 3rd ed., §16.5 (quoted; Aug 2026 draft). Graves et al. (2006), Proc. ICML, 369–376.
DEFINITION
Encoder-decoder with attention

The decoder is “a conditional language model that attends to the encoder representation”. Course gloss: at each output step, attention turns scores into a probability distribution over encoder frames and takes the weighted sum: a soft, learned alignment.

J&M §16.3.2 (quoted). Chan et al. (2016), Proc. ICASSP, 4960–4964; Vaswani et al. (2017), NeurIPS 30, 5998–6008.
DEFINITION · EXTRA
Self-supervised pretraining

“This pretraining phase doesn’t require transcripts; just unlabeled speech files.” Course gloss: hide parts of the input and train the model to identify what was hidden; then fine-tune on a small labelled set.

J&M §16.4 (quoted; its example is HuBERT). The lecture's example: Baevski et al. (2020), NeurIPS 33, 12449–12460 (wav2vec 2.0).
DEFINITION · EXTRA
Weak supervision at scale

Course wording: training on very large quantities of found audio-transcript pairs of mixed quality, trading label cleanliness for coverage: Whisper's 680,000 hours, 117,000 of them in 96 languages other than English.

Radford et al. (2023), Proc. ICML, PMLR 202, 28492–28518 (Whisper).
Exercises: pencils out · ~7 minutes, in pairs

Try it: collapse and align

EX 8.1

Collapse these frame strings: (a) h h ∅ e e l ∅ l o   (b) h e l l o o. One of them fails to spell “hello”: which, and why?

EX 8.2

State the blank's two jobs. Then explain why removing ∅ from the alphabet would make some English words impossible for CTC to output at all.

EX 8.3

DTW's path and an attention row are both alignments. Give two precise differences between them.

EX 8.4

Suppose base Whisper writes “we go makan lah” as “we go McCann”. Give two model-internal reasons, and name one remedy that changes the model rather than the input.

Worked answers in the appendix at the end of this deck.

Before you go

Four things to remember

1. CTC labels every frame with a letter or the blank ∅, merges repeats, deletes blanks, and trains on the sum over every path that collapses to the transcript

2. The table trick, three kinds of arithmetic: minimum (DTW), maximum (Viterbi), sum (the forward algorithm, and now CTC)

3. Attention is a soft alignment learned from data: for each output token, a weighted sum over the frames

4. With no lexicon file and no separate language model, what remains to change in a recognizer is its training data

Radar: Quiz 2 (start of Week 12) covers Weeks 6 to 11: today is week three of six. Assignment 1 is due Friday 23 October, the end of Week 10.

This week

This week's readings

REQ3Blue1Brown, “Attention in transformers, step-by-step” (26 min)

Rewatch it after today's attention slides. 3blue1brown.com/lessons/attention

REQHannun (2017), “Sequence Modeling with CTC”, Distill

Read through the collapsing rule; the animations are the point. distill.pub/2017/ctc

REQHugging Face Audio Course, Unit 3: “Transformer architectures for audio”

The CTC and Seq2Seq pages. huggingface.co/learn/audio-course/chapter3/introduction

REQJurafsky & Martin, SLP (3rd ed., free online): Ch 16, §16.3 “The Encoder-Decoder Architecture for ASR” and §16.5 “CTC”

Aug 2026 draft. Its blank ϵ is our ∅; by today's rule its §16.5.1 example [ϵ b b] collapses to b. web.stanford.edu/~jurafsky/slp3/16.pdf

OPTRadford et al. (2023), Whisper, and Baevski et al. (2020), wav2vec 2.0

The introduction and approach sections of each. arxiv.org/abs/2212.04356 · arxiv.org/abs/2006.11477

OPTXu, “Utilising ASR in Linguistic Research”, Ch 1 “Applying large pre-trained models” (online tutorial)

Whisper and wav2vec 2.0 on your own recordings; its wav2vec 2.0 code is today's greedy CTC decoding. chenzixu.rbind.io/resources/3asr/sr1

Next week

Readings for Week 9

REQKoenecke et al. (2020), “Racial disparities in automated speech recognition”, PNAS 117(14), 7684–7689

Short and open access; read the methods as a model of an audit. pnas.org/doi/10.1073/pnas.1915768117

REQChodroff (2018), “Corpus Phonetics Tutorial”, arXiv:1811.05553

The Montreal Forced Aligner sections; updated web version at eleanorchodroff.com/tutorial. arxiv.org/abs/1811.05553

REQHugging Face Audio Course, Unit 5

WER, CER and text normalization. huggingface.co/learn/audio-course/chapter5/introduction

OPTMcAuliffe et al. (2017), “Montreal Forced Aligner”, Proc. Interspeech 2017, 498–502

Readable now that you know HMMs. isca-archive.org/interspeech_2017/mcauliffe17_interspeech.html

OPTMarkl (2022), “Language variation and algorithmic bias”, FAccT ’22, 521–534

Bias beyond US English, written by a sociolinguist. dl.acm.org/doi/10.1145/3531146.3533117

ʃ

Sum the paths.
Learn the alignment.

That is end-to-end recognition, and you ran it on Singapore English.  Week 9, “The Audit”: word error rates across speakers, bias in the numbers, and forced alignment, which brings back the classical HMM, lexicon and Viterbi as a measuring tool.

Sources

References

Jurafsky, D. & Martin, J. H. Speech and Language Processing, 3rd ed. (Aug 2026 draft). Ch. 16, §16.3 (encoder-decoder), §16.4 (self-supervised models), §16.5 (CTC). Quoted on the formal slide.

Graves, A., Fernández, S., Gomez, F. & Schmidhuber, J. (2006). “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks.” Proc. ICML 2006, 369–376. CTC.

Hannun, A. (2017). “Sequence modeling with CTC.” Distill. doi:10.23915/distill.00008

Chan, W., Jaitly, N., Le, Q. & Vinyals, O. (2016). “Listen, attend and spell: a neural network for large vocabulary conversational speech recognition.” Proc. ICASSP 2016, 4960–4964. Attention for ASR.

Vaswani, A. et al. (2017). “Attention is all you need.” NeurIPS 30, 5998–6008. The transformer.

Baevski, A., Zhou, H., Mohamed, A. & Auli, M. (2020). “wav2vec 2.0: a framework for self-supervised learning of speech representations.” NeurIPS 33, 12449–12460.

Boigne, J. (2021). “An illustrated tour of wav2vec 2.0.” Blog post. jonathanbgn.com/2021/09/30/illustrated-wav2vec-2.html

Maye, J., Werker, J. F. & Gerken, L. (2002). “Infant sensitivity to distributional information can affect phonetic discrimination.” Cognition 82(3), B101–B111.

Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C. & Sutskever, I. (2023). “Robust speech recognition via large-scale weak supervision.” Proc. ICML 2023, PMLR 202, 28492–28518. Whisper.

Koenecke, A., Choi, A. S. G., Mei, K. X., Schellmann, H. & Sloane, M. (2024). “Careless Whisper: speech-to-text hallucination harms.” Proc. FAccT ’24, 1672–1681.

Pasad, A., Chou, J.-C. & Livescu, K. (2021). “Layer-wise analysis of a self-supervised speech representation model.” IEEE ASRU 2021, 914–921. Probing.

Koh, J. X. et al. (2019). “Building the Singapore English National Speech Corpus.” Proc. Interspeech 2019, 321–325. The NSC.

jensenlwt. whisper-small-singlish-122k. Hugging Face model card (accessed Sep 2026). huggingface.co/jensenlwt/whisper-small-singlish-122k. The Singlish fine-tune.

Sanderson, G. (3Blue1Brown). (2024). “Attention in transformers, step-by-step.” 3blue1brown.com/lessons/attention

Hugging Face. Audio Course, Unit 3: “Transformer architectures for audio.” huggingface.co/learn/audio-course

Appendix: worked answers

Answers

EX 8.1

(a) merge repeats: h ∅ e l ∅ l o; delete blanks: hello. (b) merge repeats: h e l o; delete blanks: helo: no blank separated the two l's, so they merged. A double letter survives only if at least one ∅ sits between the two copies (l ∅ l, l l ∅ l).

EX 8.2

Job 1: a label for frames that start no new letter (silence, transitions, the middle of a sound). Job 2: a separator that keeps real double letters apart. Without ∅, repeats always merge, so words with a double letter (“book”, “hello”, “lorry”) cannot be output, and silent frames must be labelled with some letter.

EX 8.3

First: DTW's path is hard (each frame of one recording is matched to specific frames of the other), while an attention row is soft (a probability distribution; every frame gets some weight). Second: DTW's monotonic shape is fixed by a rule, while attention's near-diagonal shape is learned from data, and nothing stops it from looking back or jumping. (Also acceptable: DTW compares two given signals; attention aligns an output being generated to its input.)

EX 8.4

Reason 1: the decoder is a strong English language model, so unclear audio becomes frequent English words and names. Reason 2: “makan” is probably rare in the training transcripts, so its word pieces score low against look-alikes. One remedy: fine-tune on Singapore English speech (for example the NSC), which adapts what the model hears (the encoder) and what it expects (the decoder); the practical's A/B is one before-and-after example.