MauryaHub PYQ Practice

ee4001_2026T1_Q2_NA.pdf

Speech Technology · Quiz 2 · Jan 2026

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 0.0 marks

**Instructions:** All **option-based questions are MCQs,** and all **fill-in-the-blank questions are short answer** **questions.** For Questions **3 (1 to 14),** answer all questions. • In **MCQs,** you must write the **correct option** and also the **complete text of that option.** • You must then provide a **detailed explanation** of why that specific option is the correct answer. • In **short answer questions,** write the answer clearly with proper explanation wherever needed. • In all **problem-solving or example-based questions,** present the solution with **all intermediate** **steps clearly shown.** • Step-wise explanation is **compulsory** for all such questions. • The **final answer** must be written clearly at the end.
  1. Instructions has been mentioned above.
  2. This Instructions is just for a reference & not for an evaluation.

A published solution is not available for this question yet.

Question 3 MCQ · 25.0 marks

1. (1 point) A key advantage of continuous HMM over discrete HMM in speech recognition is: (a) No need for feature extraction (b) Direct modeling of real-valued acoustic features (c) No training data required (d) No state transitions required 2. (1 point) Ina typical 3-state phoneme HMM, self-loops are mainly used to model: (a) Vocabulary growth (b) Speaker identity (c) Variable duration (d) Word frequency 3. (1 point) If attention weights shift backward in time during decoding, what does it violate? (a) CTC assumption (b) Vocabulary constraint (c) Encoder out (d) Monotonic alignment 4. (1 point) Fill in the blank: The predictor in RNN-T is similar to a language model, since it conditions on 5. (1 point) The fine-tuning step of wav2vec 2.0/HuBERT using a simple linear layer for ASR typically uses: (a) Connectionist Temporal Classification (CTC) loss (b) Crose-entropy loss with phoneme labels ] (d) Contrastive loss 6. (1 point) Which statement is false about CTC? (a) CTC assumes conditional independence between output labels at different time steps. (b) CTC decoding often uses greedy or beam search. (c) CTC explicitly learns alignments between input and output during training. (d) CTC introduces a special blank token. 7. (1 point) A CNN has 3 layers with strides: Layer 1: stride = 4 Layer 2: stride = 4 Layer 3: stride = 5 What is the total downsampling factor? (a) 40 (b) 60 (c) 80 (d) 100 8. (1 point) At a specific node (t,1) in the RNN-T lattice, the Joiner network receives rt” and hf”. Ifthe Joiner outputs a probability for the "blank” symbol, what does a transition to the next state represent? (a) A move to (t,u +1) (b) A move to (f+1,u +1) (c) Termination of the sequence. (d) A move to (t+1,1) 9. (1 point) In an attention-based ASR system, suppose the attention weights are uniform across all encoder time steps. What effect does this have on the alignment between input speech frames and output tokens? (a) Good alignment (b) Poor alignment (c) Faster decoding (d) No decoding 10. (1 point) Suppose GPT predicts probabilities for the next token: “cat” > 0.7, “dog” — 0.2, “car” 3 0.1 If the true token is “cat”, what is the cross-entropy loss? (a) -log(a7) (b) -log(a2) (©) “og(0.1) (A) -og(0.3) 11. (2 points) In automatic speech recognition (ASR), how does the Word Error Rate (WER) typically vary across CTC, RNN-T, and AED (Attention-based Encoder-Decoder) models, assuming similar training data and model capacity? (a) WER(CTC) < WER(RNN-T) < WER(AED) (b) WER(RNN-T) < WER(AED) < WER(CTC) (c) WER(CTC) = WER(RNN-T) = WER(AED) (a) WER(AED) < WER(RNN-T) < WER(CTC) [[IMAGE:719ea63b4a327d30_6_5]]
Source diagram or notation
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.