Skip to main content

Chapter 6 · 4 hours

Speech and Channel Coding

IOE past exam questions

Past questions and answers

20 questions set from this chapter, 5 of them more than once. Most asked first.

  • Asked 5 times
  • 2080 Baisakh · 2+6 marks
  • 2079 Bhadra · 3+5 marks
  • 2076 Bhadra · 3+5 marks
  • 2075 Bhadra · 2+6 marks
  • 2072 Magh · 2+6 marks

What are the characteristics of speech signal? Explain the operation of linear predictive coder (LPC) with neat block diagram.

Answer

Characteristics of speech signal

  • Bandwidth: most energy and intelligibility lie in about 300–3400 Hz (telephone band).
  • Non-uniform amplitude distribution: small amplitudes are much more likely than large ones (Laplacian-like pdf), so non-uniform quantization helps.
  • Short-term stationarity: speech is quasi-stationary over 10–30 ms, so it can be analysed in frames.
  • Voiced and unvoiced sounds: voiced sounds (vowels) are quasi-periodic with a pitch of roughly 80–350 Hz; unvoiced sounds (s, f, sh) are noise-like.
  • Formants: the vocal tract resonances give spectral peaks (formants) that carry most information.
  • High correlation between adjacent samples and pitch periods, so speech is predictable.
  • Silence: about 40–60% of a conversation is pauses, used by voice activity detection.
  • Non-flat spectrum: energy falls at higher frequencies.

Linear predictive coder (LPC)

LPC is a vocoder that models the vocal tract as an all-pole digital filter and sends only the filter parameters plus the excitation type, instead of the waveform. Each sample is predicted as a linear combination of the previous p samples:

ŝ(n) = Σₖ₌₁ᵖ aₖ · s(n − k)

The coefficients aₖ are chosen to minimise the mean square prediction error e(n) = s(n) − ŝ(n) over a frame, by solving the normal (Yule–Walker) equations with the Levinson–Durbin algorithm.

 Encoder (transmitter)
 speech -> [A/D] -> [frame 20 ms] -+-> [LPC analysis:
                                   |    a1..ap, gain G]
                                   +-> [pitch / V-UV
                                        detector]
            -> quantize & multiplex -> bit stream

 Decoder (receiver)
  [pulse train @ pitch]--+
                         +-[V/UV switch]-(×G)-+
  [white noise gen.]-----+                    |
                                              v
             speech <- [D/A] <- [all-pole filter 1/A(z)]

Operation

  1. Speech is sampled (8 kHz) and divided into frames of about 20 ms.
  2. LPC analysis finds about 10 predictor coefficients (vocal-tract filter) and the gain G for each frame.
  3. A pitch detector decides voiced/unvoiced and measures the pitch period.
  4. Coefficients, gain, pitch and V/UV flag are quantized and transmitted; e.g. LPC-10 sends 54 bits per 22.5 ms frame = 2.4 kbps.
  5. At the receiver, the excitation is a periodic pulse train (voiced) or white noise (unvoiced), scaled by G.
  6. The excitation drives the all-pole synthesis filter H(z) = G / (1 − Σ aₖz⁻ᵏ) to rebuild speech.

Merits: very low bit rate (2.4–4.8 kbps). Drawback: synthetic, "robotic" quality. Improved versions (multi-pulse, RELP, CELP, RPE-LTP in GSM) send a better excitation signal.

  • Asked 4 times
  • 2075 Bhadra · 3 marks
  • 2072 Asoj · 3 marks
  • 2070 Bhadra · 3 marks
  • 2070 Magh · 4 marks

Write a short note on Viterbi decoding algorithm.

Answer

The Viterbi algorithm is a maximum-likelihood decoding method for convolutional codes. It finds the path through the code trellis whose output sequence is closest to the received sequence (minimum Hamming distance for hard decision, or Euclidean distance for soft decision).

Steps

  1. Draw the trellis: 2ᴷ⁻¹ states (K = constraint length), with branches labelled by encoder output bits.
  2. For each received symbol group, compute the branch metric = Hamming distance between received bits and each branch's output.
  3. Add–compare–select: for every state, add branch metric to the path metric of each incoming path; keep the smaller one as the survivor, discard the other.
  4. Repeat for the whole sequence (the encoder is flushed to state 00 with tail zeros).
  5. Trace back the survivor ending in state 00 with the smallest metric; its input bits are the decoded data.

Example: rate 1/2, K = 3 code, g₁ = 111, g₂ = 101.

  • Data 1011 + two tail zeros → transmitted 11 10 00 01 01 11.
  • Received with one error: 11 10 10 01 01 11.
  • The survivor ending in state 00 has metric 1 and gives 1 0 1 1 0 0, so the error is corrected.

Features: complexity grows as 2ᴷ⁻¹ (linear in message length); soft-decision gives about 2 dB extra gain. Used in GSM, IS-95 and as the MLSE equalizer.

  • Asked 3 times
  • 2082 Bhadra · 4 marks
  • 2081 Baisakh · 4 marks
  • 2072 Asoj · 8 marks

What are different characteristics of speech signals? How are they used in designing of coders?

Answer

Characteristics of speech signals

  1. Limited bandwidth – intelligible speech lies mostly in 300–3400 Hz.
  2. Non-uniform probability density – low amplitudes occur far more often than high ones (Laplacian/gamma-like pdf).
  3. Non-zero autocorrelation – adjacent samples (125 µs apart at 8 kHz) are highly correlated (correlation coefficient often > 0.85).
  4. Non-flat spectrum – energy is concentrated at low frequencies, with resonant peaks called formants.
  5. Quasi-periodicity (pitch) – voiced sounds repeat with a pitch period of about 3–12 ms (80–350 Hz); unvoiced sounds are noise-like.
  6. Short-term stationarity – statistics stay nearly constant over 10–30 ms segments.
  7. Silence periods – speakers are silent about 40–60% of the time.
  8. Perceptual masking and phase insensitivity – the ear is relatively insensitive to phase and to noise near strong spectral components.

How these are used in coder design

CharacteristicUse in coder design
Bandwidth 300–3400 HzSample at 8 kHz; filter out other bands
Non-uniform pdfNon-uniform quantization (μ-law/A-law) gives more levels to small amplitudes
Sample correlationPredictive coding (DPCM, ADPCM) sends only the prediction error, fewer bits
Non-flat spectrumSub-band and adaptive transform coding give more bits to strong bands
FormantsLPC models the vocal tract as an all-pole filter; only ~10 coefficients sent
Pitch periodicityLong-term (pitch) predictor; voiced excitation = pulse train
Short-term stationarityProcess in 20 ms frames; update parameters once per frame
SilenceVoice activity detection and discontinuous transmission save power and capacity
Perceptual maskingPerceptual weighting filter in CELP shapes noise under the speech spectrum

Example: the GSM full-rate RPE-LTP coder uses 20 ms frames, an 8th-order short-term LPC filter (formants) and a long-term predictor (pitch) to reach 13 kbps instead of 64 kbps PCM.

  • Asked 2 times
  • 2081 Bhadra · 3+5 marks
  • 2073 Magh · 2+3 marks

Describe vocoders with block diagram. Briefly explain different kind of vocoders.

Answer

Vocoders

A vocoder (voice coder) is a source coder that does not try to reproduce the speech waveform. It extracts the parameters of a model of human speech production – vocal-tract filter, excitation type (voiced/unvoiced), pitch and gain – sends only these, and the receiver synthesizes speech from them. This gives very low bit rates (about 1.2–4.8 kbps) at the cost of natural quality.

 Analyzer (Tx)
 speech -+-> [spectral analysis] ----+
         +-> [pitch, V/UV detector] -+-> MUX -> channel

 Synthesizer (Rx)
 channel -> DEMUX -+-> pitch, V/UV
                   |    -> [pulse gen | noise gen]
                   |              |
                   +-> spectrum ->[synthesis filter]
                                  |
                                  v speech

The model: speech = excitation (pulse train for voiced, random noise for unvoiced) passed through a time-varying filter (vocal tract).

Kinds of vocoders

  1. Channel vocoder – A bank of band-pass filters (about 16–19 bands) measures the energy in each band; the band envelopes plus pitch/V-UV are sent. The receiver modulates a matching filter bank with the excitation. Earliest vocoder.
  2. Formant vocoder – Sends the frequencies and amplitudes of the first three or four formants (vocal-tract resonances) plus pitch. Very low rate (< 1.2 kbps) but formants are hard to track accurately.
  3. Cepstrum vocoder – Uses the cepstrum (inverse FFT of log spectrum) to separate the excitation (high quefrency) from the vocal-tract response (low quefrency); low cepstral coefficients and pitch are sent.
  4. Voice-excited vocoder – Sends a narrowband (baseband) portion of the actual speech as excitation instead of a pitch detector output, avoiding pitch-detection errors.
  5. Linear predictive coder (LPC) – Models the vocal tract as an all-pole filter with ~10 coefficients; LPC-10 runs at 2.4 kbps. Its improved forms (multi-pulse, RELP, CELP) are used in GSM and IS-95.

Example of bit rate saving: 64 kbps PCM vs 2.4 kbps LPC-10 is a reduction of about 27 times, which is why vocoder principles (in the form of hybrid coders like CELP) are used in mobile systems.

  • Asked 2 times
  • 2078 Chaitra · 2+6 marks
  • 2070 Magh · 4+4 marks

Why do we need speech coding techniques? Explain the basic concept of VOCODER.

Answer

Need for speech coding

Speech coding compresses digitized speech into fewer bits while keeping acceptable quality. It is needed because:

  • Plain 64 kbps PCM uses too much of the scarce, expensive radio spectrum; low-rate coders (e.g. 13 kbps GSM, 8 kbps CELP) let more users share a channel and raise system capacity.
  • Lower bit rate means narrower bandwidth, less transmit power and longer battery life.
  • Coders can include protection of important bits and work well with channel coding, keeping quality high over noisy, fading links.
  • It allows voice activity detection and discontinuous transmission, reducing interference.
  • Digital speech can be encrypted for privacy and multiplexed easily with data.

Basic concept of VOCODER

A vocoder is a parametric speech coder based on a model of how humans produce speech:

  • Source (excitation): the vocal cords produce a periodic pulse train for voiced sounds (pitch 80–350 Hz) or turbulent noise for unvoiced sounds.
  • Filter: the vocal tract (throat, mouth, nose) acts as a slowly varying filter whose resonances are the formants.

Instead of sending the waveform, the vocoder analyses each 10–30 ms frame and sends only:

  1. Vocal-tract filter parameters (band energies, formants or LPC coefficients),
  2. Voiced/unvoiced decision,
  3. Pitch period, and
  4. Gain (loudness).
 Tx: speech -> [analysis] -> {filter params, V/UV,
                              pitch, gain} -> encode
 Rx: pitch pulses --+
                    +--[V/UV]--(×gain)--[vocal tract
     noise ---------+                    filter]--> speech

The receiver uses a pulse generator or noise generator as the excitation and drives a synthesis filter set by the received parameters.

Waveform coder vs vocoder

PointWaveform coderVocoder
What is sentSamples of waveformModel parameters
Bit rate16–64 kbps1.2–4.8 kbps
QualityNaturalSynthetic
ExamplePCM, ADPCMChannel, LPC-10

Result: rates of 1.2–4.8 kbps, good intelligibility but synthetic quality. Types include channel, formant, cepstrum, voice-excited and linear predictive vocoders.

  • 2082 Bhadra · 2+3 marks

Define turbo coding. Construct 8 by 8 Hadamard code.

Answer

Turbo coding

Turbo coding is a powerful forward error-correction scheme that uses two (or more) recursive systematic convolutional (RSC) encoders in parallel, separated by an interleaver, and is decoded iteratively by two soft-in soft-out decoders that exchange information. It performs within about 0.5–1 dB of the Shannon limit and is used in 3G (WCDMA, CDMA2000) and 4G LTE.

 d --+---------------------> systematic bits
     +-->[RSC encoder 1]---> parity 1
     +-->[interleaver]->[RSC encoder 2]-> parity 2

8 × 8 Hadamard code

A Hadamard (Walsh) matrix is built recursively (Sylvester construction), where H′ is the complement of H:

 H1 = [0]

 H2 = | H1  H1 |  = | 0 0 |
      | H1  H1'|    | 0 1 |

 H4 = | H2  H2 |  = | 0 0 0 0 |
      | H2  H2'|    | 0 1 0 1 |
                    | 0 0 1 1 |
                    | 0 1 1 0 |

 H8 = | H4  H4 |      (H' = complement of H)
      | H4  H4'|

Writing out H8:

         c1 c2 c3 c4 c5 c6 c7 c8
 W0:      0  0  0  0  0  0  0  0
 W1:      0  1  0  1  0  1  0  1
 W2:      0  0  1  1  0  0  1  1
 W3:      0  1  1  0  0  1  1  0
 W4:      0  0  0  0  1  1  1  1
 W5:      0  1  0  1  1  0  1  0
 W6:      0  0  1  1  1  1  0  0
 W7:      0  1  1  0  1  0  0  1

Each row is one 8-bit Walsh codeword. Any two rows differ in exactly 4 positions (in ±1 form, with 0 → +1 and 1 → −1, their inner product is 0), so the rows are orthogonal, as required for CDMA spreading codes.

  • 2082 Baisakh · 4+4 marks

What are the characteristics of speech signals? Explain the basic working principle of formant and channel vocoder.

Answer

Characteristics of speech signals

  • Bandwidth mainly 300–3400 Hz.
  • Non-uniform amplitude pdf: small amplitudes are most probable.
  • High correlation between adjacent samples (predictable).
  • Non-flat spectrum with resonant peaks called formants.
  • Voiced sounds are quasi-periodic (pitch 80–350 Hz); unvoiced sounds are noise-like.
  • Quasi-stationary over 10–30 ms frames.
  • About 40–60% silence in conversation.

Channel vocoder

The channel vocoder is the oldest vocoder (Dudley, 1939). It represents the short-time spectrum by the energy in a set of frequency bands.

 Tx: speech-+-[BPF1]-[rectify+LPF]-> env 1 -+
            +-[BPF2]-[rectify+LPF]-> env 2 -+-> MUX
            +-[BPFn]-[rectify+LPF]-> env n -+
            +-[pitch & V/UV detector]-------+
 Rx: [pulse/noise]-+-(×env1)-[BPF1]-+
                   +-(×env2)-[BPF2]-+-(Σ)-> speech
                   +-(×envn)-[BPFn]-+
  1. Speech passes through a bank of about 16–19 band-pass filters covering 200–3200 Hz.
  2. Each filter output is rectified and low-pass filtered to give its slowly varying spectral envelope.
  3. Envelopes are sampled (about every 20 ms) and sent with pitch and voiced/unvoiced information.
  4. The receiver drives an identical filter bank with a pulse train (voiced) or noise (unvoiced), scales each band by its envelope and sums the bands.

Formant vocoder

The formant vocoder uses the fact that the speech spectrum is described mainly by its first few formants.

  1. A formant tracker estimates the frequencies (and amplitudes) of the first three or four formants in each frame.
  2. These values, plus pitch and V/UV decision, are transmitted.
  3. The receiver uses resonators tuned to the received formant frequencies, excited by pulses or noise, to synthesize speech.

It needs fewer parameters than a channel vocoder, so the bit rate is very low (often below 1.2 kbps), but formants are hard to track accurately, especially in noise, so it is less used in practice.

PointChannel vocoderFormant vocoder
Spectrum sent asBand energies (16–19)3–4 formant frequencies
Bit rate~2.4 kbps< 1.2 kbps
Main difficultyMany channelsAccurate formant tracking
  • 2081 Baisakh · 6 marks

With the help of a block diagram, explain the operation of a vocoder.

Answer

A vocoder is a speech coder that analyses speech into the parameters of a speech-production model and transmits only those parameters. The receiver synthesizes speech from them, giving bit rates of about 1.2–4.8 kbps.

Speech model: excitation source (pulse train at the pitch frequency for voiced sounds, random noise for unvoiced sounds) → time-varying vocal-tract filter → speech.

 ANALYZER (transmitter)
 speech->[frame]-+->[vocal tract analysis]-> a_k
                 +->[V/UV decision]      -> flag
                 +->[pitch estimator]    -> T0
                 +->[energy measure]     -> G
           a_k, flag, T0, G -> quantize -> MUX

 SYNTHESIZER (receiver)
 [pulse gen @ T0]--+
                   +-[V/UV]-(×G)->[vocal tract
 [noise gen]-------+               filter a_k]->speech

Operation

  1. Framing: sampled speech is cut into short frames (10–30 ms) in which it is nearly stationary.
  2. Spectral analysis: the vocal-tract response is estimated, as band energies (channel vocoder), formant frequencies (formant vocoder) or predictor coefficients (LPC vocoder).
  3. Excitation analysis: a detector decides whether the frame is voiced or unvoiced and measures the pitch period.
  4. Gain: the frame energy is computed.
  5. These parameters are quantized, multiplexed and transmitted (e.g. LPC-10: 2.4 kbps).
  6. Synthesis: at the receiver the V/UV flag selects either the pulse generator (at the received pitch) or the noise generator. The excitation is scaled by the gain and passed through a filter set by the received spectral parameters to recreate speech.

Merits: very low bit rate, good intelligibility, suits secure and narrow channels. Demerits: synthetic quality, sensitive to pitch errors and background noise.

  • 2080 Bhadra · 6 marks

Explain sub-band coding with transmitter and receiver.

Answer

Sub-band coding (SBC) is a frequency-domain speech coding method in which the speech band is split into several frequency sub-bands by a filter bank, and each sub-band is coded separately with a number of bits matched to its perceptual importance and energy.

Idea: speech energy and its importance to the ear are not uniform over 0–4 kHz. Low bands (with pitch and the first formants) get more bits; high bands get fewer. Quantization noise is also kept within each band, so it is masked by the speech in that band.

 TRANSMITTER
         +-[BPF1]-[↓D]-[ADPCM, n1 bits]-+
 s(n) ---+-[BPF2]-[↓D]-[ADPCM, n2 bits]-+-[MUX]-> channel
         +-[BPFk]-[↓D]-[ADPCM, nk bits]-+

 RECEIVER
            +-[decode1]-[↑D]-[BPF1]-+
 ch->[DEMUX]+-[decode2]-[↑D]-[BPF2]-+-(Σ)-> s^(n)
            +-[decodek]-[↑D]-[BPFk]-+
   (↓D: decimate, ↑D: interpolate)

Transmitter

  1. Speech sampled at 8 kHz is passed through a bank of k band-pass filters (often quadrature mirror filters, QMF), e.g. 4–8 bands.
  2. Each band output is decimated (down-sampled) to its Nyquist rate, so the total sample rate stays 8 kHz.
  3. Each sub-band is encoded, usually with ADPCM, using a bit allocation nᵢ based on band energy and perception (e.g. 4–5 bits in low bands, 2 bits in high bands).
  4. Coded bits of all bands are multiplexed and transmitted.

Receiver

  1. The bit stream is demultiplexed into the sub-band streams.
  2. Each stream is decoded, interpolated (up-sampled) and passed through the matching synthesis filter.
  3. The band outputs are added to reconstruct the speech. QMF filters cancel the aliasing created by decimation.

Advantages: noise in each band is controlled separately; bits are allocated to what the ear hears; good quality at 16–32 kbps (ITU G.722 wideband audio uses two-band SBC with ADPCM at 64 kbps). Disadvantage: filter bank adds delay and complexity.

  • 2080 Bhadra · 4 marks

Construct Hadamard matrix for H8.

Answer

A Hadamard matrix of order N is a square matrix whose rows are mutually orthogonal. It is constructed recursively by the Sylvester method:

H₂ₙ = [ Hₙ Hₙ ; Hₙ −Hₙ ] (in ±1 form), or with Hₙ′ = complement of Hₙ (in 0/1 form).

Step 1: H₁ = [+1]

Step 2:

 H2 = | +1 +1 |
      | +1 -1 |

Step 3:

 H4 = | H2  H2 | = | +1 +1 +1 +1 |
      | H2 -H2 |   | +1 -1 +1 -1 |
                   | +1 +1 -1 -1 |
                   | +1 -1 -1 +1 |

Step 4:

 H8 = | H4  H4 |
      | H4 -H4 |

      | +1 +1 +1 +1 +1 +1 +1 +1 |
      | +1 -1 +1 -1 +1 -1 +1 -1 |
      | +1 +1 -1 -1 +1 +1 -1 -1 |
      | +1 -1 -1 +1 +1 -1 -1 +1 |
      | +1 +1 +1 +1 -1 -1 -1 -1 |
      | +1 -1 +1 -1 -1 +1 -1 +1 |
      | +1 +1 -1 -1 -1 -1 +1 +1 |
      | +1 -1 -1 +1 -1 +1 +1 -1 |

In binary form (+1 → 0, −1 → 1):

         c1 c2 c3 c4 c5 c6 c7 c8
 W0:      0  0  0  0  0  0  0  0
 W1:      0  1  0  1  0  1  0  1
 W2:      0  0  1  1  0  0  1  1
 W3:      0  1  1  0  0  1  1  0
 W4:      0  0  0  0  1  1  1  1
 W5:      0  1  0  1  1  0  1  0
 W6:      0  0  1  1  1  1  0  0
 W7:      0  1  1  0  1  0  0  1

Check: the inner product of any two different rows is 0 (e.g. row 2 · row 3 = 1 − 1 − 1 + 1 + 1 − 1 − 1 + 1 = 0), and H8·H8ᵀ = 8·I. These 8 rows are the Walsh codes used for channelization in CDMA.

  • 2080 Chaitra · 3+5 marks

Explain the characteristics of speech. Explain 7-bit Hamming code method with relevant example.

Answer

Characteristics of speech

  • Most energy and intelligibility in 300–3400 Hz.
  • Non-uniform amplitude pdf: small amplitudes are more likely.
  • Strong correlation between adjacent samples, so speech is predictable.
  • Non-flat spectrum with formant peaks.
  • Voiced sounds are quasi-periodic (pitch 80–350 Hz); unvoiced sounds are noise-like.
  • Quasi-stationary over 10–30 ms; about 40–60% silence in conversation.

7-bit Hamming code

The (7,4) Hamming code is a linear block code that adds 3 parity bits to 4 data bits. Minimum distance d_min = 3, so it corrects any single-bit error (or detects two). Parity bits sit at positions that are powers of 2:

Position1234567
Bitp1p2d1p3d2d3d4

Even-parity equations (each parity covers the positions whose binary index has that bit set):

 p1 = d1 ⊕ d2 ⊕ d4   (positions 1,3,5,7)
 p2 = d1 ⊕ d3 ⊕ d4   (positions 2,3,6,7)
 p3 = d2 ⊕ d3 ⊕ d4   (positions 4,5,6,7)

Example – encoding: data 1011 → d1 = 1, d2 = 0, d3 = 1, d4 = 1.

 p1 = 1 ⊕ 0 ⊕ 1 = 0
 p2 = 1 ⊕ 1 ⊕ 1 = 1
 p3 = 0 ⊕ 1 ⊕ 1 = 0
 Codeword (p1 p2 d1 p3 d2 d3 d4) = 0 1 1 0 0 1 1

Example – error correction: suppose bit 5 is flipped; received word = 0 1 1 0 1 1 1.

 s1 = r1 ⊕ r3 ⊕ r5 ⊕ r7 = 0⊕1⊕1⊕1 = 1
 s2 = r2 ⊕ r3 ⊕ r6 ⊕ r7 = 1⊕1⊕1⊕1 = 0
 s3 = r4 ⊕ r5 ⊕ r6 ⊕ r7 = 0⊕1⊕1⊕1 = 1
 Syndrome (s3 s2 s1) = 101 = 5

The syndrome points to position 5, so that bit is inverted back to 0, giving 0110011 and data 1011. A syndrome of 000 means no error.

Code rate = 4/7 ≈ 0.57. Such block codes protect the most important speech-coder bits in mobile systems.

  • 2079 Chaitra · 3+5 marks

Write the steps to satisfy condition for Hadamard code. Construct 8 by 8 Hadamard code that satisfies above conditions.

Answer

Conditions (steps) for a Hadamard code

A Hadamard (Walsh) code is a set of N codewords of length N (N = 2ᵏ) taken from the rows of a Hadamard matrix. The conditions it must satisfy are:

  1. Square matrix: N rows and N columns, entries only 0/1 (or +1/−1).
  2. First row (and first column) all zeros (all +1), i.e. the matrix is normalized.
  3. Equal weight: every row except the first has exactly N/2 zeros and N/2 ones.
  4. Orthogonality: any two different rows agree in exactly N/2 places and differ in N/2 places. In ±1 form, the inner product of two different rows is 0, so H·Hᵀ = N·I.

Steps to construct (Sylvester method) – these steps guarantee the above conditions:

  1. Start with H₁ = [0].
  2. Form H₂ₙ from Hₙ by placing Hₙ in the top-left, top-right and bottom-left blocks and its complement Hₙ′ in the bottom-right block:
 H2n = | Hn  Hn  |
       | Hn  Hn' |
  1. Repeat until the required size N is reached (1 → 2 → 4 → 8).

Construction of 8 × 8 Hadamard code

 H1 = [0]

 H2 = | H1  H1 |  = | 0 0 |
      | H1  H1'|    | 0 1 |

 H4 = | H2  H2 |  = | 0 0 0 0 |
      | H2  H2'|    | 0 1 0 1 |
                    | 0 0 1 1 |
                    | 0 1 1 0 |

 H8 = | H4  H4 |      (H' = complement of H)
      | H4  H4'|

Writing H8 in full:

         c1 c2 c3 c4 c5 c6 c7 c8
 W0:      0  0  0  0  0  0  0  0
 W1:      0  1  0  1  0  1  0  1
 W2:      0  0  1  1  0  0  1  1
 W3:      0  1  1  0  0  1  1  0
 W4:      0  0  0  0  1  1  1  1
 W5:      0  1  0  1  1  0  1  0
 W6:      0  0  1  1  1  1  0  0
 W7:      0  1  1  0  1  0  0  1

Verification:

  • Row W0 is all zeros; every other row has four 0s and four 1s.
  • Example: W1 = 01010101 and W2 = 00110011 agree in positions 1, 4, 5, 8 and differ in 2, 3, 6, 7, i.e. 4 agreements and 4 disagreements, so they are orthogonal. The same holds for every pair (checked for all 28 pairs).
  • Minimum distance between codewords = N/2 = 4.

These 8 orthogonal codewords are used as Walsh spreading codes to separate channels in CDMA (IS-95 uses the 64 × 64 version).

  • 2077 Chaitra · 6 marks

Discuss the different characteristics of speech signal.

Answer

Speech signals have special statistical and spectral properties. Speech coders use these to cut the bit rate below 64 kbps PCM.

  1. Limited bandwidth – Speech covers about 50 Hz–10 kHz, but almost all intelligibility lies in 300–3400 Hz, so telephone speech is sampled at 8 kHz.
  2. Non-uniform amplitude distribution – The pdf of speech amplitude is peaked at zero (Laplacian/gamma-like): small amplitudes are much more frequent than large ones. This leads to non-uniform (μ-law, A-law) and adaptive quantization.
  3. High sample-to-sample correlation – At 8 kHz sampling, adjacent samples are strongly correlated (coefficient often above 0.85). The future sample can be predicted from past ones, which is used by DPCM, ADPCM and LPC.
  4. Non-flat (non-white) spectrum – Energy falls with frequency, and there are resonance peaks called formants due to the vocal tract. Sub-band and transform coders exploit this by allocating bits to strong bands.
  5. Voiced and unvoiced sounds – Voiced sounds (vowels) are quasi-periodic, produced by vibrating vocal cords; the repetition rate is the pitch (about 80–160 Hz for men, 160–350 Hz for women). Unvoiced sounds (s, f, sh) are random, noise-like. Vocoders use a pulse or noise excitation accordingly.
  6. Short-term stationarity – The vocal tract changes slowly, so speech is nearly stationary over 10–30 ms. Coders process speech in frames of about 20 ms.
  7. Long-term periodicity – Successive pitch periods look similar; long-term (pitch) predictors remove this redundancy.
  8. Silence/pauses – In a conversation each speaker is active only about 40% of the time. Voice activity detection and discontinuous transmission use this to save power and reduce interference.
  9. Perceptual properties – The ear is fairly insensitive to phase and to noise masked by nearby strong components; perceptual weighting in CELP uses this.
  • 2074 Bhadra · 6 marks

Describe the operation of any two source coders used in speech coding.

Answer

Speech source coders remove redundancy in speech to reduce the bit rate. Two widely used ones are described below.

1. Adaptive Differential PCM (ADPCM) – a waveform coder

ADPCM uses the high correlation between speech samples. It predicts each sample from past samples and quantizes only the prediction error, using a quantizer whose step size adapts to the signal level.

 s(n)->(+)-e(n)->[adaptive quantizer]-+-> e_q(n) out
        ^ -                  ^         |
        |             [step adapt]     v
        +-- s^(n) <--[predictor]<-----(+)<-- s^(n)
               s~(n) = s^(n) + e_q(n) (reconstructed)
  • The predictor forms ŝ(n) from previous reconstructed samples.
  • Error e(n) = s(n) − ŝ(n) has a much smaller range than s(n), so fewer bits are needed.
  • Step size and predictor coefficients adapt to the changing speech level.
  • The decoder contains the same predictor, adding e_q(n) to ŝ(n).
  • ITU G.726 ADPCM gives toll quality at 32 kbps (4 bits/sample) – half of 64 kbps PCM. Used in DECT.

2. Linear Predictive Coder (LPC) – a vocoder

LPC models speech as an excitation passed through an all-pole vocal-tract filter, and sends only model parameters.

 Tx: speech->[20 ms frames]->[LPC analysis a1..a10, G]
                           ->[pitch, V/UV] -> encode
 Rx: [pulses | noise]->(×G)->[1/A(z) filter]->speech
  • Each frame: about 10 predictor coefficients aₖ (from minimising prediction error, ŝ(n) = Σ aₖs(n−k)), gain G, pitch and voiced/unvoiced flag.
  • The receiver excites the filter 1/A(z) with a pulse train (voiced) or noise (unvoiced).
  • LPC-10 works at 2.4 kbps; quality is synthetic. Hybrid versions (CELP, RPE-LTP) send a better excitation and are used in mobile systems.
  • 2074 Bhadra · 5 marks

Write a short note on convolutional encoding and decoding.

Answer

Convolutional coding is a forward error-correction method in which each group of k input bits produces n output bits that depend on the current input and the previous (K − 1) inputs stored in a shift register. It is described by (n, k, K), with code rate r = k/n and constraint length K.

Encoding – example (2, 1, 3) encoder, rate 1/2, generators g₁ = 111, g₂ = 101:

 input ->[ m ]->[ s1 ]->[ s2 ]
          |       |       |
          +--(⊕)--+--(⊕)--+---> v1 = m ⊕ s1 ⊕ s2
          |               |
          +------(⊕)------+---> v2 = m ⊕ s2
          output: v1 v2 (alternately)

Input 1011 followed by two flushing zeros gives output 11 10 00 01 01 11. The encoder can be shown as a state diagram, a tree, or a trellis (states = 2ᴷ⁻¹ = 4).

Decoding – the Viterbi algorithm (maximum likelihood) is standard:

  1. For each received bit pair, compute branch metrics (Hamming distance) on the trellis.
  2. At each state, add-compare-select: keep only the survivor path with the smallest total metric.
  3. At the end, trace back from state 00 to get the decoded bits.

Example: received 11 10 10 01 01 11 (one error) is decoded back to 1011, as the best path has metric 1. Soft-decision decoding gives about 2 dB extra gain. Sequential decoding is an alternative for long K.

Used in GSM (rate 1/2, K = 5), IS-95 (K = 9) and satellite links.

  • 2074 Magh · 4+2+2 marks

Describe waveform and voice coding techniques. Mention characteristics of speech. Draw a suitable diagram of GSM CODEC.

Answer

Waveform coding and voice (source) coding techniques

Waveform coders try to reproduce the time waveform of speech sample by sample. They work for any signal (speech, music, tones), are robust, but need higher bit rates (16–64 kbps).

  • Time domain: PCM (64 kbps, μ/A-law companding), DPCM, ADPCM (32 kbps), delta modulation and CVSD.
  • Frequency domain: sub-band coding and adaptive transform coding, which give more bits to the important bands.

Voice coders (vocoders / source coders) do not copy the waveform. They model speech production (excitation + vocal tract filter) and send only model parameters: spectral shape, pitch, voiced/unvoiced flag and gain. Bit rates are very low (1.2–4.8 kbps) but quality is synthetic.

  • Types: channel, formant, cepstrum, voice-excited and linear predictive (LPC) vocoders.

Hybrid coders (multi-pulse LPC, RELP, CELP, RPE-LTP) combine both: LPC model plus a coded excitation chosen by analysis-by-synthesis, giving good quality at 4–16 kbps. These are used in cellular systems.

PointWaveform coderVocoder
SendsWaveform samplesModel parameters
Bit rate16–64 kbps1.2–4.8 kbps
QualityHigh, naturalSynthetic
ComplexityLowHigh

Characteristics of speech

  • Band 300–3400 Hz; non-uniform amplitude pdf; strong sample correlation.
  • Non-flat spectrum with formants; voiced (pitch 80–350 Hz) and unvoiced sounds.
  • Quasi-stationary over 10–30 ms; about 40–60% silence.

GSM CODEC

GSM full-rate uses the RPE-LTP (Regular Pulse Excited – Long Term Prediction) coder at 13 kbps (260 bits per 20 ms frame).

 speech 8 kHz, 13-bit
   |
 [Pre-processing] (offset removal, pre-emphasis)
   |
 [STP: LPC analysis, 8 coefficients] --> LAR (36 bits)
   |
 [Short-term analysis filter] -> residual
   |
 [LTP: pitch lag & gain] --------------> (4 x 9 bits)
   |
 [RPE grid selection & coding] --------> (4 x 47 bits)
   |
 MUX: 36 + 36 + 188 = 260 bits / 20 ms = 13 kbps

The decoder reverses this: RPE decoding → LTP synthesis → short-term LPC synthesis filter → post-processing → speech.

  • 2074 Magh · 5 marks

Write a short note on turbo coding.

Answer

Turbo coding is a forward error-correction (FEC) technique, introduced by Berrou, Glavieux and Thitimajshima in 1993, that uses parallel concatenated recursive systematic convolutional (RSC) codes joined by an interleaver, with iterative soft decoding. It comes within about 0.5–1 dB of the Shannon limit.

Encoder

 d(k) --+------------------------------> x (systematic)
        +-->[RSC encoder 1]------------> y1 (parity 1)
        +-->[interleaver]->[RSC enc. 2]-> y2 (parity 2)
                 |
        [puncturing / MUX] -> rate 1/3 (or 1/2)
  • The data bits are sent unchanged (systematic part).
  • Encoder 1 produces parity from the data in natural order; encoder 2 produces parity from an interleaved (permuted) copy.
  • The interleaver makes the two parity streams nearly independent, so a pattern that is weak for one encoder is strong for the other.
  • Puncturing can raise the code rate from 1/3 to 1/2.

Decoder

     +--------------------------------------+
     v                                      |
 x,y1->[SISO decoder 1]->[interleave]->[SISO decoder 2]
                                  y2-->^        |
           extrinsic info  <-[de-interleave]<---+
                                   after N iterations
                                   -> hard decision
  • Two soft-input soft-output decoders (MAP/BCJR or SOVA) exchange extrinsic information (log-likelihood ratios).
  • Each iteration improves the reliability of the bit estimates; 4–8 iterations are typical.

Merits: near-Shannon-limit performance at low SNR. Demerits: decoding delay and complexity, and an "error floor" at very low BER. Used in 3G (WCDMA, CDMA2000), LTE data channels and deep-space links.

  • 2071 Bhadra · 4+4 marks

Explain the operation of formant vocoder. What are the characteristics of speech signal?

Answer

Formant vocoder

The formant vocoder is based on the fact that the short-time spectrum of speech is mainly described by its formants, the resonant frequencies of the vocal tract. Normally the first three or four formants (F1 ≈ 300–800 Hz, F2 ≈ 800–2500 Hz, F3 ≈ 2–3 kHz) carry most of the information.

 Tx: speech-+->[formant tracker]-> F1,F2,F3 (+ amplitudes)
            +->[pitch & V/UV detector] -> pitch, flag
            +->[energy] -> gain        -> MUX -> channel

 Rx: [pulse gen | noise gen] (by V/UV, pitch)
            |
            v (×gain)
     [resonator F1]->[resonator F2]->[resonator F3]
            |
            v speech

Operation

  1. Speech is divided into frames of 10–30 ms.
  2. A formant tracker finds the frequencies (and sometimes bandwidths/amplitudes) of the first 3–4 formants in each frame, e.g. by peak-picking the spectrum or from LPC poles.
  3. A pitch detector decides voiced/unvoiced and measures pitch.
  4. Formants, pitch, V/UV flag and gain are coded and transmitted.
  5. The receiver selects a pulse train (voiced) or noise (unvoiced), scales it and passes it through resonators tuned to the received formant frequencies, which rebuild the spectral envelope.

Merits: very few parameters, so very low bit rate (below about 1.2 kbps). Demerits: formants are hard to track reliably (they merge or are weak, especially in noise), so quality is poor; hence formant vocoders are rarely used commercially.

Characteristics of speech signal

  • Main energy in 300–3400 Hz.
  • Non-uniform amplitude pdf (small values most likely).
  • High correlation between successive samples.
  • Non-flat spectrum with formant peaks.
  • Voiced sounds periodic (pitch 80–350 Hz), unvoiced sounds noise-like.
  • Quasi-stationary over 10–30 ms.
  • Silence for about 40–60% of the time in a conversation.
  • 2071 Magh · 2+6 marks

What is channel coding? Explain types of linear predictive coder.

Answer

Channel coding

Channel coding is the process of adding controlled redundancy to the data before transmission so that the receiver can detect and correct errors caused by noise, fading and interference. It lowers the BER for a given SNR (coding gain) at the cost of extra bandwidth or lower data rate. Types: block codes (Hamming, BCH, Reed–Solomon), convolutional codes (Viterbi decoding) and turbo/LDPC codes. In mobile systems it is used with interleaving.

Types of linear predictive coder

All LPC coders model the vocal tract with an all-pole filter 1/A(z), whose ~10 coefficients are found by linear prediction, ŝ(n) = Σ aₖ s(n − k). They differ in how the excitation is produced.

1. Basic LPC vocoder (LPC-10)

  • Excitation is a pulse train (voiced) or white noise (unvoiced) chosen by a V/UV decision and pitch detector.
  • 2.4 kbps; intelligible but synthetic speech.

2. Multi-pulse excited LPC (MPE-LPC)

  • Excitation is a small number of pulses (e.g. 4–8 per 5 ms) whose positions and amplitudes are found by analysis-by-synthesis: the encoder tries pulses and keeps those that minimise the perceptually weighted error.
  • No V/UV decision needed; good quality at about 9.6 kbps. GSM RPE-LTP uses regularly spaced pulses (regular pulse excitation).

3. Code-excited LPC (CELP)

  • Excitation is chosen from a codebook of stored random (Gaussian) sequences. For each sub-frame the encoder passes every codeword through the synthesis filter (with long-term pitch predictor) and picks the one with least weighted error; only its index and gain are sent.
 [codebook]->(×gain)->[pitch filter]->[1/A(z)]-> s^(n)
                                                  |
 s(n) ---------------------------------(-)<-------+
                    error -> [weighting filter] -> min?
  • Good quality at 4.8–8 kbps; used in IS-95 (QCELP), FS-1016 and ACELP in GSM EFR.

4. Residual excited LPC (RELP)

  • The prediction residual e(n) = s(n) − ŝ(n) is low-pass filtered, down-sampled and sent as excitation; the receiver regenerates the high-frequency part by spectral folding.
  • No pitch detection needed; good quality at about 9.6 kbps.
  • 2070 Bhadra · 2+6 marks

What is vocoder? Explain any two predictive coders.

Answer

Vocoder

A vocoder (voice coder) is a speech source coder that analyses speech into the parameters of a speech-production model – vocal-tract filter, excitation type (voiced/unvoiced), pitch and gain – and transmits only these. The receiver synthesizes speech from them. Bit rates are very low (1.2–4.8 kbps) but quality is synthetic. Examples: channel, formant, cepstrum and LPC vocoders.

Predictive coders

Predictive coders use linear prediction: each sample is estimated from past samples, ŝ(n) = Σₖ₌₁ᵖ aₖ s(n − k), and the vocal tract is modelled by the all-pole filter 1/A(z). Two widely used predictive coders:

1. Multi-pulse excited LPC (MPE-LPC)

 [pulse generator]->[1/A(z)]-> s^(n) --(-)<-- s(n)
        ^                               |
        +--[error minimisation]<-[weighting filter]
  • LPC coefficients are computed for each frame as usual.
  • Instead of a single pitch pulse or noise, the excitation has several pulses per sub-frame (e.g. 4–8 per 5 ms).
  • Pulse positions and amplitudes are found one at a time by analysis-by-synthesis, minimising the perceptually weighted error between original and synthesized speech.
  • No voiced/unvoiced decision or pitch detection errors; good quality at about 9.6 kbps. GSM's RPE-LTP is a variant with regularly spaced pulses.

2. Code-excited LPC (CELP)

 [codebook index i]->(×G)->[pitch predictor]->[1/A(z)]
                                                 |
                       s(n) --(-)<--- s^(n) -----+
                               |
                [perceptual weighting] -> choose i, G
  • A codebook of, say, 1024 stored random excitation vectors is kept at both ends.
  • For each 5 ms sub-frame the encoder tries the codevectors through a long-term (pitch) predictor and short-term LPC filter, and selects the one giving minimum weighted error.
  • Only the codebook index, gain, pitch parameters and LPC coefficients are sent.
  • High quality at 4.8–8 kbps; used in IS-95 (QCELP), FS-1016 (4.8 kbps) and ACELP (GSM EFR, AMR).

(A third type, residual excited LPC (RELP), sends the low-pass filtered prediction residual as excitation.)

Questions from Old Question Collection (EX 751 and BEI EX 715) (IOE exam papers: EX 751 (BEX) 2070 Bhadra to 2080 Chaitra and EX 715 (BEI) 2079 Bhadra to 2082 Bhadra). Answers are written for this site; check them against your class notes.

Chapter titles and hours from the IOE syllabus ↗