Blanks and Attention.
Chenzi Xu · chenzi.xu@ntu.edu.sg
1. The fundamental equation of classical ASR?
Ŵ = argmaxW P(X | W) · P(W): the acoustic likelihood times the language-model prior
2. DTW and Viterbi fill the same kind of table. What differs in the cells?
DTW keeps the minimum cost of the cells leading in; Viterbi keeps the maximum probability
3. What does a softmax output layer produce?
one probability per class: positive numbers that sum to one
4. How does a network learn its weights?
the loss (cross-entropy) scores its mistakes; gradient descent moves each weight a step downhill
5. In the hybrid recognizer, what did the network replace, and what stayed?
it replaced the Gaussian emission scores; the HMM, lexicon, Viterbi and n-gram language model stayed
Today one network replaces every part of echo 1.
Released: 25 Sep (Week 7)
Due: Fri 23 Oct (Week 10)
Weight: 30%
Everything task-specific is in the brief on NTULearn.
In the practical: write the CTC collapse, then test Whisper on Singlish speech.
End-to-end recognition: train one network on (audio, text) pairs; it learns the acoustic model, the spelling and the word sequences together.
Do not choose one alignment: add up the probabilities of all of them.
At each output step, compute a weight for every frame: a soft alignment, in which every frame counts a little.
Two hundred frames. Two letters.
Don't choose. Sum.
Connectionist temporal classification (CTC): the network labels every frame,
a collapse rule turns the labels into a transcript, and training adds up the probabilities of every labelling that collapses to the correct one.
Spot the difference: b o o k collapses to bok; b o ∅ o k collapses to book. Without a blank between them, repeats always merge. The ∅ needs no silence: in b b o o ∅ o k it sits inside the vowel.
n n ∅ o o → n ∅ o → no
∅ n o o ∅ → ∅ n o ∅ → no
n ∅ n o → n ∅ n o → nno
many frame strings spell “no”; one of these does not
k k ∅ æ æ æ ∅ t t → kæt
k ∅ æ æ ∅ ∅ t → kæt
∅ k æ ∅ t t ∅ → kæt
phone labels here; English CTC models usually output letters
Greedy decoding can miss the most probable transcript; beam search (Week 6) often finds it. See the take-home stretch.
| transcript | paths | total P |
| … | … | … |
| … | … | … |
| … | … | … |
| … | … | … |
…and 105 more: all 109 transcripts add up to 1; “no” wins on probability, not on number of paths.
Week 5, DTW: the minimum over paths.
Week 6, Viterbi: the maximum; the forward algorithm: the sum.
Week 8, CTC: the forward sum; its −log is the training loss.
J&M §16.5.2.
* only to skip a ∅ between two different letters. Never three rows up: that would jump over a letter.
Each frame's label is chosen without looking at the labels on other frames. J&M §16.5.1: “CTC makes a strong conditional independence assumption”
CTC spells by sound: a real CTC model wrote “makan” as MACKIN. Systems add a separate language model at decoding time, or use an attention decoder
Labels follow the audio in order: suitable for speech, not for translation, where word order changes
U letters need at least U frames, plus a ∅ between double letters: “book” needs 5 (b o ∅ o k)
CTC: every alignment, summed.
Next: alignment, learned.
After the break: attention, then wav2vec 2.0 and Whisper.
score each encoder vector against the decoder's state (a dot product) → softmax into weights → weighted sum
Transformer: a stack of blocks; each block is self-attention, then feedforward layers. The 3Blue1Brown attention video shows one step by step.
Simplified: here each letter has its own row of scores for this one clip. A real network learns the scoring itself (the matrices behind the dot products) from thousands of clips, so it can align clips it has never seen.
Learns from unlabelled audio first, then from transcribed speech (1 hour can be enough), with CTC
An encoder-decoder with attention, trained on 680,000 hours of web audio: the model in today's practical
Pretrained on 53,000 unlabelled hours, then fine-tuned: 1 hour of labels beat the previous best, which had 100 hours of labels; 10 minutes of labels still gave 4.8% WER on clean read English (with a language model). Good news for languages with few transcripts.
Baevski et al. (2020), NeurIPS 33, 12449–12460. J&M §16.4 describes HuBERT, a close relative.
To learn more: Boigne (2021), “An Illustrated Tour of Wav2vec 2.0”. jonathanbgn.com/2021/09/30/illustrated-wav2vec-2.html
A the transcript, word for word, with the particles and “makan”
B something fluent but anglicised: “Can, let's go to McCann”
C nothing: the audio is not in its training data
jensenlwt, whisper-small-singlish-122k (Hugging Face model card); Radford et al. (2023), §6; Koenecke et al. (2024).
Every part can be inspected · a bad pronunciation is fixed by editing one lexicon line · phonetic knowledge entered each module directly
Lower error rates, given enough training data · coarticulation modelled, not assumed away · new languages without hand-built lexicons · one model to deploy
The linguist's role now:
Run Whisper.
Then find its errors.
Open the Week 8 notebook (link on NTULearn). In the next hour: the CTC collapse, whisper-small on three clips, and a base versus fine-tuned comparison on one Singapore English clip.
Stretch (take-home): a two-frame grid with P(a) = 0.4 and P(∅) = 0.6 on both frames. Greedy reads ∅ ∅ and outputs nothing, yet the three paths that spell “a” add up to 0.64. Explain why a search over transcripts picks “a”. Illustrative numbers.
illustrative layout: one clip, two models, two transcripts, two WERs: your first audit table
The stretch is take-home; stages 1 to 4 are the in-class core.
A collapse function you finished, two models' transcripts of two clips, a WER table, and a record of what Whisper does with silence.
If a model download stalls, the notebook uses stored transcripts of the provided clips.
The intuition of CTC is “to output a single character for every frame of the input”; a function then “collapses all repeated letters and then removes all blanks”. Course gloss: P(transcript | audio) is the sum over every frame-label path that collapses to it.
Jurafsky & Martin, SLP 3rd ed., §16.5 (quoted; Aug 2026 draft). Graves et al. (2006), Proc. ICML, 369–376.The decoder is “a conditional language model that attends to the encoder representation”. Course gloss: at each output step, attention turns scores into a probability distribution over encoder frames and takes the weighted sum: a soft, learned alignment.
J&M §16.3.2 (quoted). Chan et al. (2016), Proc. ICASSP, 4960–4964; Vaswani et al. (2017), NeurIPS 30, 5998–6008.“This pretraining phase doesn’t require transcripts; just unlabeled speech files.” Course gloss: hide parts of the input and train the model to identify what was hidden; then fine-tune on a small labelled set.
J&M §16.4 (quoted; its example is HuBERT). The lecture's example: Baevski et al. (2020), NeurIPS 33, 12449–12460 (wav2vec 2.0).Course wording: training on very large quantities of found audio-transcript pairs of mixed quality, trading label cleanliness for coverage: Whisper's 680,000 hours, 117,000 of them in 96 languages other than English.
Radford et al. (2023), Proc. ICML, PMLR 202, 28492–28518 (Whisper).Collapse these frame strings: (a) h h ∅ e e l ∅ l o (b) h e l l o o. One of them fails to spell “hello”: which, and why?
State the blank's two jobs. Then explain why removing ∅ from the alphabet would make some English words impossible for CTC to output at all.
DTW's path and an attention row are both alignments. Give two precise differences between them.
Suppose base Whisper writes “we go makan lah” as “we go McCann”. Give two model-internal reasons, and name one remedy that changes the model rather than the input.
Worked answers in the appendix at the end of this deck.
1. CTC labels every frame with a letter or the blank ∅, merges repeats, deletes blanks, and trains on the sum over every path that collapses to the transcript
2. The table trick, three kinds of arithmetic: minimum (DTW), maximum (Viterbi), sum (the forward algorithm, and now CTC)
3. Attention is a soft alignment learned from data: for each output token, a weighted sum over the frames
4. With no lexicon file and no separate language model, what remains to change in a recognizer is its training data
Radar: Quiz 2 (start of Week 12) covers Weeks 6 to 11: today is week three of six. Assignment 1 is due Friday 23 October, the end of Week 10.
Rewatch it after today's attention slides. 3blue1brown.com/lessons/attention
Read through the collapsing rule; the animations are the point. distill.pub/2017/ctc
The CTC and Seq2Seq pages. huggingface.co/learn/audio-course/chapter3/introduction
Aug 2026 draft. Its blank ϵ is our ∅; by today's rule its §16.5.1 example [ϵ b b] collapses to b. web.stanford.edu/~jurafsky/slp3/16.pdf
The introduction and approach sections of each. arxiv.org/abs/2212.04356 · arxiv.org/abs/2006.11477
Whisper and wav2vec 2.0 on your own recordings; its wav2vec 2.0 code is today's greedy CTC decoding. chenzixu.rbind.io/resources/3asr/sr1
Short and open access; read the methods as a model of an audit. pnas.org/doi/10.1073/pnas.1915768117
The Montreal Forced Aligner sections; updated web version at eleanorchodroff.com/tutorial. arxiv.org/abs/1811.05553
WER, CER and text normalization. huggingface.co/learn/audio-course/chapter5/introduction
Readable now that you know HMMs. isca-archive.org/interspeech_2017/mcauliffe17_interspeech.html
Bias beyond US English, written by a sociolinguist. dl.acm.org/doi/10.1145/3531146.3533117
Sum the paths.
Learn the alignment.
That is end-to-end recognition, and you ran it on Singapore English. Week 9, “The Audit”: word error rates across speakers, bias in the numbers, and forced alignment, which brings back the classical HMM, lexicon and Viterbi as a measuring tool.
Jurafsky, D. & Martin, J. H. Speech and Language Processing, 3rd ed. (Aug 2026 draft). Ch. 16, §16.3 (encoder-decoder), §16.4 (self-supervised models), §16.5 (CTC). Quoted on the formal slide.
Graves, A., Fernández, S., Gomez, F. & Schmidhuber, J. (2006). “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks.” Proc. ICML 2006, 369–376. CTC.
Hannun, A. (2017). “Sequence modeling with CTC.” Distill. doi:10.23915/distill.00008
Chan, W., Jaitly, N., Le, Q. & Vinyals, O. (2016). “Listen, attend and spell: a neural network for large vocabulary conversational speech recognition.” Proc. ICASSP 2016, 4960–4964. Attention for ASR.
Vaswani, A. et al. (2017). “Attention is all you need.” NeurIPS 30, 5998–6008. The transformer.
Baevski, A., Zhou, H., Mohamed, A. & Auli, M. (2020). “wav2vec 2.0: a framework for self-supervised learning of speech representations.” NeurIPS 33, 12449–12460.
Boigne, J. (2021). “An illustrated tour of wav2vec 2.0.” Blog post. jonathanbgn.com/2021/09/30/illustrated-wav2vec-2.html
Maye, J., Werker, J. F. & Gerken, L. (2002). “Infant sensitivity to distributional information can affect phonetic discrimination.” Cognition 82(3), B101–B111.
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C. & Sutskever, I. (2023). “Robust speech recognition via large-scale weak supervision.” Proc. ICML 2023, PMLR 202, 28492–28518. Whisper.
Koenecke, A., Choi, A. S. G., Mei, K. X., Schellmann, H. & Sloane, M. (2024). “Careless Whisper: speech-to-text hallucination harms.” Proc. FAccT ’24, 1672–1681.
Pasad, A., Chou, J.-C. & Livescu, K. (2021). “Layer-wise analysis of a self-supervised speech representation model.” IEEE ASRU 2021, 914–921. Probing.
Koh, J. X. et al. (2019). “Building the Singapore English National Speech Corpus.” Proc. Interspeech 2019, 321–325. The NSC.
jensenlwt. whisper-small-singlish-122k. Hugging Face model card (accessed Sep 2026). huggingface.co/jensenlwt/whisper-small-singlish-122k. The Singlish fine-tune.
Sanderson, G. (3Blue1Brown). (2024). “Attention in transformers, step-by-step.” 3blue1brown.com/lessons/attention
Hugging Face. Audio Course, Unit 3: “Transformer architectures for audio.” huggingface.co/learn/audio-course
(a) merge repeats: h ∅ e l ∅ l o; delete blanks: hello. (b) merge repeats: h e l o; delete blanks: helo: no blank separated the two l's, so they merged. A double letter survives only if at least one ∅ sits between the two copies (l ∅ l, l l ∅ l).
Job 1: a label for frames that start no new letter (silence, transitions, the middle of a sound). Job 2: a separator that keeps real double letters apart. Without ∅, repeats always merge, so words with a double letter (“book”, “hello”, “lorry”) cannot be output, and silent frames must be labelled with some letter.
First: DTW's path is hard (each frame of one recording is matched to specific frames of the other), while an attention row is soft (a probability distribution; every frame gets some weight). Second: DTW's monotonic shape is fixed by a rule, while attention's near-diagonal shape is learned from data, and nothing stops it from looking back or jumping. (Also acceptable: DTW compares two given signals; attention aligns an output being generated to its input.)
Reason 1: the decoder is a strong English language model, so unclear audio becomes frequent English words and names. Reason 2: “makan” is probably rare in the training transcripts, so its word pieces score low against look-alikes. One remedy: fine-tune on Singapore English speech (for example the NSC), which adapts what the model hears (the encoder) and what it expects (the decoder); the practical's A/B is one before-and-after example.