HG4052 · Speech Synthesis & Recognition
i
HG4052 · Speech Synthesis & Recognition / Week 4

Thirteen Honest
Numbers.

Waveforms and Features.

Chenzi Xu · chenzi.xu@ntu.edu.sg

Δ
Today

Roadmap

In the practical hour you build the entire MFCC pipeline yourself, one stage per notebook cell.

t
The shape of the morning

Today, by the clock

09:40  Lecture I: frames, the ear's warp, and Week 1's matrix at work
10:40  Break (10 min)
10:50  Lecture II: the cepstrum, pitch, and what we discard
11:40  Break + Praat and Colab setup (10 min)
11:50  Practical: build the MFCC pipeline, stage by stage
12:50  Wrap-up: what to remember, what comes next
Before new math · 3 minutes, tell your neighbour

Five echoes from Week 3

1. To capture a sound with energy up to 8 kHz, sample at least how fast?

16 kHz: Nyquist wants twice the highest frequency present

2. Source or filter: which one decides which vowel you hear?

the filter: the vocal tract, whose resonances are the formants

3. One column of a spectrogram comes from doing what to a short slice?

a Fourier transform: dot the slice with a probe sine at every frequency: the spectrum is a row of dot products

4. Why fade a slice's edges before analysing it?

a hard cut is a click, and a click smears energy across every frequency

5. Vertical striations in a wideband spectrogram? Horizontal ones in a narrowband?

each vertical striation is one glottal pulse; each horizontal striation is one harmonic of F0

The bridge from Week 3

Automated narrow transcription, of the acoustics

You read spectrograms; a recognizer takes numbers. The MFCCs (mel frequency cepstral coefficients; Davis & Mermelstein 1980) are the classic form of those numbers: about 13 per centisecond, capturing which sound is happening: the standard recognizer front end from 1980 into the 2010s. Here is the assembly line, today's whole lecture on one slide:

waveformWeek 3's list of numbers
→
frames~25 ms slices, every 10 ms
→
windowfade the edges, no clicks
→
spectrum“STFT”: what Praat does per column
→
mel filterbankbend the axis like an ear
→
logloudness is logarithmic
→
DCTkeep hills, drop pickets
→
13 MFCCsper frame

STFT (short-time Fourier transform) = “compute a spectrum every 10 ms”. Window = “fade each slice in and out”: Week 3’s Hamming taper, guarding against the click a hard cut would add.

Short-time analysis

Slicing speech

Articulators are always moving, and the signal is only “steady” for ~10–30 ms at a time. So: 25 ms frames, one every 10 ms (engineers call the step the hop): 100 frames per second.

The window-size trade-off: 25 ms sits comfortably inside a vowel but smears a 15 ms stop burst. Narrowband vs wideband, now as an engineering decision.

Short-time analysis · continued

From frames to columns

zeros (librosa)a mirror image (torch)t = 0, the recording starts12.5 ms of paddingframe 0: half padding, half soundreads 25 ms (17.5 to 42.5)its column: 10 ms wide, same centre−1001020304050607080msthe wavethe framesthe columns

The frame is 25 ms of sound, the window the taper laid on it, the column the stripe of picture its spectrum becomes. Each frame is stamped at the centre of the 25 ms it read, and its column is painted 10 ms wide at that same centre: frames overlap in what they read, never in what they paint. Before t = 0 there is nothing to read, so the tools pad 12.5 ms of invented signal (zeros in librosa, a mirror image in torch) and frame 0 sits at 0; Praat pads nothing and centres its first frame half a window in.

Short-time analysis · the price of the ruler

Same tile area, your choice of shape

One tile = the footprint of one measurement: the patch of time × frequency that a single spectrogram value answers for. Whatever falls inside one tile is reported as one blur.

25 ms window · tiles 25 ms wide × 40 Hz tall (1/25 ms)
15 ms burst2 harmonics, 120 Hz apart0100 ms1 kHz0✓ harmonics land in separate rows✗ the burst hides inside one 25 ms tileone tile: 25 ms × 40 Hz
5 ms window · tiles 5 ms wide × 200 Hz tall (1/5 ms)
15 ms burst2 harmonics, 120 Hz apart0100 ms1 kHz0✗ both harmonics share one 200 Hz tile✓ the burst is pinned to its 5 ms tilesone tile: 5 ms × 200 Hz

A window of length T cannot say when inside those T ms something happened, nor tell apart two frequencies closer than about 1/T, so every tile has the same area: choosing T only chooses its shape. Anything briefer than the tile (the 15 ms burst in a 25 ms window) is reported by every frame that touches it: smeared across ~40 ms of columns.

Short-time analysis · what one column holds

FFT of one frame gives 257 numbers

1 · 512 samples of probe room

  • 25 ms = 400 samples; zero-padded to 512 (HTK, Kaldi), or a 512-sample window outright (librosa’s default)
  • 512 ÷ 16,000 per second = 32 ms

2 · Probes: whole cycles per frame

  • the FFT pretends the frame repeats, so a probe must join up with itself
  • k cycles in 32 ms = k ÷ 0.032 s = k × 31.25 Hz

3 · Bins 0 to 256: 257 numbers

  • probe 256 = 8 kHz, two samples per cycle: the fastest 16 kHz can hold (Nyquist)
  • probes 257 to 511 repeat 255 down to 1: dropped, so N/2 + 1

FFT length sets the grid (31.25 Hz); window length sets the blur (about 40 Hz for 25 ms, previous slide).

Short-time analysis · the probe trick, watched

How a probe scores the frame

1 · One dot product per bin

  • frame × probe, sample by sample, then add up (Week 1’s dot product, Week 3’s trick)
  • wiggles in step: products mostly one sign, big score; out of step: they cancel, score near 0

2 · No exact match needed

  • 700 Hz fits 22.4 cycles: probes 22 and 23 both score high, the rest get fading leftovers (leakage)
  • the window keeps the leftovers small; a sine twin of each probe catches waves that start shifted

3 · One hump per component

  • about 1/T wide (the tile); the bins only sample it, which is why a harmonic spans two or three stems on the filter slide
  • a curve through the top three bins reads 700 Hz off a 31.25 Hz grid
ɛ
Bending the axis

The ear’s ruler: the mel scale

The mel filterbank: ~26 triangular band-pass filters, spaced as the ear hears

Equal steps in mel ≈ equal steps in heard pitch (Stevens, Volkmann & Newman 1937): m = 1127 ln(1 + f / 700), near-linear below 1 kHz, logarithmic above.

Mel enters only through the filterbank: triangles evenly spaced in mel, so crowded below 1 kHz (F1, and the start of F2) and sparse above. It decides where resolution is spent.

Then a second log, on the other axis, of each filter’s energy, because loudness is logarithmic (decibels).

Result: the log mel spectrogram, the input to ASR systems such as Whisper.

Bending the axis · one triangle at a time

One filter, up close

Filter 8 of 26 spans 645 to 922 Hz. Nine FFT bins (31.25 Hz apart: 16 kHz, 512 points) fall under it.

Each bin’s power × the triangle’s height at that bin; add the nine products: one number, mel energy 8.

Edge bins count for little, the peak bin fully. Where two triangles overlap, one’s falling slope and the other’s rising slope add to exactly 1, so a bin’s energy is split between neighbours, never lost or doubled.

This weighted sum is row 8 of the matrix on the next slide: 257 slots, 248 of them zero.

A promise from Week 1, kept

The filterbank is a matrix

mel_energies = filterbank @ spectrum

Each row of the matrix is one triangle; multiplying says: “apply the same weighted-sum recipe to every row.” One line of numpy, and you'll type it in the practical.

Hold this thought: a matrix × a vector, then a squash: that is one layer of a neural network.

What bending the axis buys

Same vowel, two rulers

Linear frequency axis (0–8 kHz)

the formants (the vowel’s identity) squashed into the bottom

Mel axis: same sound

vowel detail spread out; the rarely-decisive top compressed

The mel warp spends resolution where phonetic contrast actually lives: the machine's version of a phonetician's trained attention.

m
10:40 · Break

Ten minutes.
Then the hardest idea of the week.

So far: frames, the mel warp, and the filterbank as a matrix. After the break, one trick pulls the source-filter mix apart again, with nothing but cosines.

★ The hardest idea this week

Can arithmetic un-mix the buzz from the tract?

Week 3: source × filter = the vowel. This week's trick: take the mixed signal apart again,
because “which vowel?” lives in the filter, and the buzz is mostly in the way.

ə
The cepstrum, without fear

A picket fence in front of rolling hills

log spectrum = slow hills + fast picket fence

Hills: the vocal-tract envelope, the formants: which vowel. Pickets: harmonics and F0: who, at what pitch.

The DCT (discrete cosine transform), “a spectrum of the log spectrum”, sorts slow from fast. Keep the first ~13 coefficients = trace the hills, drop the pickets. Those 13 numbers are the MFCCs.

The cepstrum as a second transform

A spectrum of a spectrum

why “cepstrum”?

spectrum → cepstrum
frequency → quefrency
filtering → liftering

The cepstrum · why the peak sits where it does

Why the peak lands at the pitch period

1 · Read the log spectrum as a waveform. Its x-axis is Hz, and it repeats: one picket every F0 Hz. At 100 Hz, ten pickets per kilohertz.

2 · Reuse Week 3’s rule. Whatever repeats every P gets a spike at 1/P. A pulse every T seconds gave harmonics every 1/T Hz; a picket every F0 Hz gives a spike at 1/F0 seconds.

3 · And 1/F0 is the pitch period. Count the pickets in one kilohertz: that number is the period in ms, and it is where the peak lands. 100 Hz → 10 ms; 200 Hz → 5 ms.

The second transform, watched live

The cepstrum, drawn

Press play: a cosine probe sweeps the log spectrum, its wiggle-rate climbing. Each rate scores one cepstral value. When the probe's rate finally matches the harmonic ripple, a single tall peak springs up: that is F0.

Two halves fall out of one sweep: low quefrency = the envelope = your 13 MFCCs, and the peak = the pitch period at quefrency = 1/F0. The MFCC path mel-smooths the ripple away first, which is why MFCCs forget the pitch.

The black box, opened: Week 3 returns

The DCT is the probe trick, again

Week 3: the spectrum is the signal dotted with sine probes that count Hz. Now run the very same trick on the log-mel spectrum, with cosine probes that count wiggles.

Each MFCC = one dot product. Low-wiggle probes trace the slow formant hills → big coefficients, worth keeping. High-wiggle probes chase the harmonic pickets → discard.

Keep the first ~13 and rebuild: the envelope returns, the pitch ripple is gone. That is the cepstrum: no magic, just dot products with cosines.

?
Commit before you see it · hands up

Say the same [ɑ] twice: once low, once an octave higher

The 13 MFCCs of the two recordings will be:

A  very different: the pitch changed, so the numbers change

B  nearly the same: it is still the same vowel

C  identical, down to the last decimal

Pick one out loud. The reason is the whole point of the cepstrum.

Why this is exactly what a phoneme classifier wants

Change the pitch, and the hills don't move

[ɑ] at F0 = 100 Hz
Same [ɑ], F0 = 200 Hz, an octave up

Different pickets, same hills → same MFCCs. The vowel survives a change of speaker pitch, which is precisely what “recognize the phoneme, ignore the voice” requires. (And the discarded pickets? They're where speaker identity hides: Week 12 will want them back.)

What the machine actually hears

39 numbers per centisecond

13 MFCCs: the spectral envelope (c0 ≈ overall energy)
+ 13 deltas: how fast each one is changing
+ 13 delta-deltas: acceleration

Deltas are the machine's formant transitions: the same coarticulation cues you read off spectrogram edges to find a stop's place of articulation. Velocity and acceleration of the spectrum, 100 times a second.

From Week 5 onward, this vector is what the course’s classic recognizers “hear”: DTW will measure distances between these, HMM states will score them, neural nets will take the closely related 80-dimensional log mel vectors.

What the machine actually hears · continued

A delta is a slope

1 · Fitted within a delta window

  • at frame t, fit a line through the window t−2 … t+2 (five frames, 50 ms); its slope is Δc[t]
  • Δc[t] = [1·(c[t+1] − c[t−1]) + 2·(c[t+2] − c[t−2])] / 10

2 · Centred on the frame

  • two frames before, two after: the change through t, not from the last frame
  • positive = rising into the coming frames; zero = steady

3 · Again, for acceleration

  • the same fit on the Δ track gives ΔΔ
  • ends are padded by repeating the first and last frame, so the first two deltas are unreliable

librosa’s delta is the same idea with a nine-frame window.

f
The one number we kept aside

Tracking F0: a signal that matches itself

Autocorrelation: slide a copy of the signal along itself, multiply the overlapping samples, and add: a dot product between the signal and its own shifted copy. A periodic sound scores low at most lags but jumps back up when the shift equals one whole period.

That peak lag is the period T, in samples. Read the pitch straight off it: F0 = fs / T.

This is the heart of Praat’s pitch tracker (newer versions low-pass the signal first, then run the same idea).

Compute it with me · pencils out

Autocorrelation by hand

THE RECIPE

r(τ) = Σ x[n] · x[n+τ]: shift the list by τ, multiply the overlapping samples, add. Try a toy signal with period 3:

x = [ 3, 1, −2, 3, 1, −2, 3, 1, −2 ]

r(0) = 3² + 1² + (−2)² + … = 42
r(1) = 3·1 + 1·(−2) + (−2)·3 + … = −9
r(2) = 3·(−2) + 1·3 + (−2)·1 + … = −16
r(3) = 3·3 + 1·1 + (−2)·(−2) + … = 28  (back up)

The first strong peak after lag 0 sits at τ = 3: that is the period, in samples. It is smaller than r(0) because fewer samples overlap, which is exactly why Praat and YIN normalize.

r(τ): the peak marks the period

Read off F0: divide the sample rate by the peak lag. F0 = fs / τ

Your recorded /ɑ/ at 16 kHz peaks near lag 80 → 16,000 / 80 = 200 Hz.

Tracking F0 · two settings that decide the answer

The window and the step

1 · Window: how well each estimate is made

  • as the window grows, the score at T climbs past T/2 and the reported F0 drops from 250 to 125 Hz
  • Praat: window = 3 periods of the pitch floor (75 Hz → 40 ms), long enough for every lag it searches

2 · Too short, too long

  • short: little overlap after the slide, so the score at T is small and T/2 can score higher (2F0)
  • long: pitch changes inside the window blur the peak, and 2T scores nearly as well as T (F0/2)

3 · Step: how often it is made

  • Praat: step = 0.75 / floor = 10 ms, so windows overlap heavily
  • it changes nothing in a single estimate, but a long step skips the turning points of a glide

Floor and ceiling set all of it: window = 3 / floor, step = 0.75 / floor, lags searched from 1 / ceiling to 1 / floor.

Tracking F0 · the same voice, product by product

Choosing the window size

Not the only way, and not always the best

Three ways to find the pitch

TIME DOMAIN

Autocorrelation

Slide the signal against itself; the first strong peak sits at one period. Praat’s classic method (“To Pitch (ac)”, Boersma 1993; newer default: a filtered variant). Intuitive, but prone to octave errors: it can lock onto 2T or T/2.

FREQUENCY DOMAIN

The cepstrum

Remember the pickets we discarded for the MFCCs? Their even spacing is F0. The cepstrum turns that regular comb into a single peak at the pitch-period quefrency: the same transform that gives you MFCCs hands you F0 in its high end (Noll 1967).

THE MODERN DEFAULT

YIN and pYIN

A difference function instead of a product, normalized to suppress octave errors (YIN, de Cheveigné & Kawahara 2002); pYIN adds probabilistic smoothing across frames (Mauch & Dixon 2014). This is librosa's pyin, the one you race against Praat.

All three agree on clean modal voice. They diverge on creak, breathy voice, and octave jumps: exactly where the phonetics lives, and exactly what the practical's F0 shoot-out is about.

A closing provocation: carry it to Week 9

We just threw away the pitch. Whose language was that safe for?

For English ASR, discarding F0 is a feature: the phonemes don't need it.

But in Cantonese, pitch is the lexicon:

詩 si1 ‘poem’  ·  史 si2 ‘history’  ·  時 si4 ‘time’

Same segments. The contrast lives entirely in the pickets we discarded. Likewise Yoruba, Thai, Vietnamese, Hausa…

THE PATTERN TO WATCH

A “neutral” engineering default that quietly assumes one language type. Tone-language systems add F0 back, when their builders think of it.

Week 9 is an audit of exactly this kind of decision.

F
11:40 · Break, then hands on keyboards

Set up Praat and Colab.
Then build the thing.

Open the Week 4 notebook, mount your Drive, and have Praat ready to record. In the next hour the pipeline on the roadmap becomes code you wrote.

The practical · 11:50–12:50

Build the pipeline, stage by stage

  • Record /i a u/ in Praat (fallbacks provided); measure one glottal period by hand, then check 1/T against Praat's pitch track
  • Frame + window: run the provided demo, describe the difference in one sentence
  • Build the mel front end: plot the triangles, then the matrix moment: filterbank @ spectrum
  • MFCCs by hand (log + DCT, one TODO), checked against librosa.feature.mfcc, and watch the pickets vanish from the envelope overlay
  • Vowels in feature space: the class's pooled /i a u/ tokens, MFCC c2 × c3: do the clusters match the vowel chart? Then cosine similarity between one [i] and everything else: Week 1’s tool on Week 4’s features.

Stretch: the F0 shoot-out: pyin vs Praat; diagnose one disagreement phonetically. Homework: mel-band resynthesis (80 → 20 → 8): which phonetic detail disappears first?

You leave with this, made by you

the class's vowels clustering in MFCC space: the chart, rediscovered again

The practical hour

The practical hour, stage by stage

1.  Record /i a u/; measure one glottal period by hand (10 min)
2.  Frame + window the vowel; hear the click when you don't (10 min)
3.  Mel front end: the triangles, then filterbank @ spectrum (15 min)
4.  Log + DCT: your MFCCs vs librosa; watch the pickets vanish (15 min)
5.  Load the provided pooled-vowels CSV; plot c2 × c3 (10 min)

The F0 shoot-out and the mel-band resynthesis are the take-home; the five boxes above are the in-class core.

YOU LEAVE WITH

A working MFCC front end you wrote yourself, one line of numpy for the filterbank, and a scatter of the class's /i a u/ in feature space that looks like the vowel chart.

If the DCT cell is unfinished, a fallback vector keeps the final plot working.

In the textbook's words

This week, formally

DEFINITION
Mel scale

“A mel is a unit of pitch defined such that pairs of sounds which are perceptually equidistant in pitch are separated by an equal number of mels.”  m = 1127 · ln(1 + f ⁄ 700).

Stevens et al. (1937); Jurafsky & Martin, SLP 3rd ed., ch. 15 (quoted, Aug 2026 draft).
DEFINITION
Cepstrum

The spectrum of the log spectrum (the name reverses the first letters of “spectrum”). It separates the slow-varying envelope (the formants) from the fast harmonic detail.

Jurafsky & Martin, SLP 3rd ed., ch. 15.
DEFINITION
MFCC

“The MFCC, mel frequency cepstral coefficients, is a useful representation of the waveform that emphasizes aspects of the signal that are relevant for detection of phonetic units”: commonly a 39-dim feature vector (J&M: 12 cepstral + 1 energy, each with Δ and ΔΔ; toolkits fold energy in as c0 of 13).

Jurafsky & Martin, SLP 3rd ed., §15.6 (quoted; Aug 2026 draft).
DEFINITION
Feature vector

The fixed-length list of numbers extracted from each short frame of speech that a recognizer takes as its input: here, one MFCC vector roughly every 10 ms.

Standard; Jurafsky & Martin, SLP 3rd ed., ch. 15–16.
Exercises: pencils out · ~6 minutes, in pairs

Try it: the ear's arithmetic

EX 4.1

Using m = 1127 · ln(1 + f⁄700): compute the mel value of 700 Hz and of 1400 Hz (take m(0) = 0). Which spans more mels: the step 0→700 Hz, or 700→1400 Hz? What does that say about where the filterbank spends its resolution?

EX 4.2

Frames are 25 ms long with a 10 ms hop. (a) How many frames per second? (b) Roughly how many frames does a 0.8-second word produce?

EX 4.3

Two recordings of [ɑ] by one speaker: F0 = 100 Hz and F0 = 150 Hz. Which part of the MFCC vector barely changes, and why is that exactly what a phoneme classifier wants?

EX 4.4

The DCT scores the log spectrum against cosine probes: a few-wiggle probe gives a big coefficient, a many-wiggle probe gives ≈ 0. (a) Which probe tracks the formant hills? (b) Keep only the high-wiggle coefficients: what survives, and is it any use to a phoneme recognizer?

Worked answers in the appendix at the end of this deck.

Before you go

Four things to carry out the door

1. A recognizer hears about 13 numbers per centisecond, not a picture.

2. The mel filterbank is a matrix; one multiply, and it is one squash away from a neural-network layer (Week 7).

3. The DCT is Week 3's probe trick again: keep the hills, drop the pickets.

4. Throwing away F0 is safe for English, and a bug for tone languages.

Quiz radar: Quiz 1 (start of Week 6) covers Weeks 1 to 5, so this week is on it. The surest revision is the notebook you just built: same pipeline, new numbers.

Before next week

This week’s readings

REQHugging Face Audio Course, Unit 1: “Introduction to audio data”

Sampling, waveform, spectrum, spectrogram, mel spectrogram, with the exact librosa calls from the notebook. huggingface.co/learn/audio-course/chapter1/audio_data

REQFayek (2016), “Speech Processing for Machine Learning: Filter banks, MFCCs and What's In-Between”

The classic step-by-step walkthrough of exactly the pipeline we build. Skim the code, focus on the figures. haythamfayek.com

REQJurafsky & Martin, SLP (3rd ed., free online): Ch 15 §15.5–15.6 “Log Mel Spectrum” & “MFCC” + Ch 16 §16.1 “The ASR Task”

The “why”, from the recognizer's point of view, and a preview of Week 5 (chapter numbers follow the Aug 2026 draft). web.stanford.edu/~jurafsky/slp3/15.pdf · /16.pdf

OPTLyons, “Mel Frequency Cepstral Coefficient (MFCC) tutorial”, PracticalCryptography.com

A second, more numerical pass over the same pipeline, useful if Fayek goes too fast. (Site is flaky; archived copy linked from NTULearn.)

OPT3Blue1Brown, “But what is the Fourier Transform?” (rewatch)

If the STFT step still feels like magic after the practical. 3blue1brown.com/lessons/fourier-transforms

Looking one week ahead

Next week’s readings

REQJurafsky & Martin, SLP Ch 16 §16.1 + the WER/evaluation section

The recognition task, and how recognizers are scored. (Chapter numbers follow the Aug 2026 draft.) web.stanford.edu/~jurafsky/slp3/16.pdf

REQMüller & Zalkow, FMP Notebooks, C3S2 “Dynamic Time Warping”

Prose and figures; the math is optional. audiolabs-erlangen.de/resources/MIR/FMP/C3/C3S2_DTWbasic.html

OPTSakoe & Chiba (1978), IEEE TASSP 26(1), 43–49

The original DTW paper, surprisingly readable now.

OPTHugging Face Audio Course, Unit 5: “Evaluation metrics for ASR”

WER hands-on, ahead of Week 5’s practical. huggingface.co/learn/audio-course/chapter5/evaluation

ʌ

Keep the hills.
Drop the pickets.

Thirteen honest numbers per centisecond: that's what a machine hears.  Next week, “The Same Word Twice”: the features meet their first recognizer: template matching and dynamic time warping.

Sources

References

Jurafsky, D. & Martin, J. H. Speech and Language Processing, 3rd ed. (Aug 2026 draft), Ch 15–16. The log mel and MFCC pipeline (§15.5–15.6); definitions quoted on the formal slide.

Stevens, S. S., Volkmann, J. & Newman, E. B. (1937). “A scale for the measurement of the psychological magnitude pitch.” JASA 8(3), 185–190. The mel scale.

Davis, S. B. & Mermelstein, P. (1980). “Comparison of parametric representations for monosyllabic word recognition.” IEEE TASSP 28(4), 357–366. The MFCC.

Noll, A. M. (1967). “Cepstrum pitch determination.” JASA 41(2), 293–309.

Boersma, P. (1993). “Accurate short-term analysis of the fundamental frequency...” IFA Proceedings 17, 97–110. Praat’s autocorrelation tracker.

de Cheveigné, A. & Kawahara, H. (2002). “YIN, a fundamental frequency estimator for speech and music.” JASA 111(4), 1917–1930.

Mauch, M. & Dixon, S. (2014). “pYIN: a fundamental frequency estimator using probabilistic threshold distributions.” ICASSP 2014.

Appendix: worked answers

Answers

EX 4.1

m(700) = 1127 · ln 2 ≈ 781 mel; m(1400) = 1127 · ln 3 ≈ 1238 mel. Step 0→700 Hz = 781 mel; step 700→1400 Hz = 1238 − 781 = 457 mel. The lower octave occupies more mels, so the filterbank packs detail where F1/F2 live and economizes higher up.

EX 4.2

(a) one frame per 10 ms hop → 100 frames/second. (b) ≈ 0.8 s ÷ 10 ms = ~80 frames (counting full frames, ⌊(800 − 25)/10⌋ + 1 = 78).

EX 4.3

The low cepstral coefficients, the spectral envelope, barely move: formants are set by vocal-tract shape, not F0. Raising the pitch re-spaces the harmonic “pickets” but leaves the “hills” put. A phoneme classifier wants vowel identity without the speaker's pitch: exactly what the MFCC keeps.

EX 4.4

(a) The few-wiggle (low-order) probe: it follows the slow envelope, i.e. the formant hills. (b) You would keep only the fast harmonic ripple (the pitch/source detail) and discard the envelope. That throws away which vowel and keeps who is speaking: useless for phoneme recognition, but exactly the information voice-identity and cloning will want back in Week 12.