MauryaHub PYQ Practice

ee4001_2026T1_ET_FN.pdf

Speech Technology · End Term · Jan 2026 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 0.0 marks

**Instructions :** All **option-based questions are MCQs,** and all **fill-in-the-blank questions are short answer** **questions.** For Questions **3 to 20,** answer all questions. • In **MCQs,** you must write the **correct option** and also the **complete text of that option.** • You must then provide a **detailed explanation** of why that specific option is the correct answer. • In **short answer questions,** write the answer clearly with proper explanation wherever needed. • In all **problem-solving or example-based questions,** present the solution with **all intermediate** **steps clearly shown.** • Step-wise explanation is **compulsory** for all such questions. • The **final answer** must be written clearly at the end.
  1. Instructions has been mentioned above.
  2. This Instructions is just for a reference & not for an evaluation.

A published solution is not available for this question yet.

Question 3 MCQ · 1.0 marks

If a speech signal is sampled at 16,000 samples per second, how many samples are contained in a 10-millisecond frame?
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 4 MCQ · 1.0 marks

How does HuBERT differ from Wav2Vec 2.0 in target generation? (a) Uses labeled phonemes (b) Uses online quantization only. (c) Uses offline clustering (e.g., k-means). (d) Does not use masking.
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 5 MCQ · 1.0 marks

Why is ”deduplication” a common step when tokenizing speech signals into phonetic tokens? (a) To remove noise from the original audio recording. (b) To account for the fact that speech units (like vowels) exist for multiple time frames (e.g., 70- 100ms) and map to the same cluster ID. (c) To increase the vocabulary size of the LLM. (d) To convert continuous SSL embeddings into discrete vectors.
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 6 MCQ · 1.0 marks

In a neural codec framework using Residual Vector Quantization (RVQ),how many codewords are typically generated per vector? (a) Number of RVQ layers (b) Size of each codebook (c) Number of encoder layers (d) Sampling rate of the audio
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 7 MCQ · 1.0 marks

You are training a Speech LLM. Your audio codec outputs 4 RVQ codebooks.The frame rate is 50 Hz. If you choose a strictly flattened Autoregressive (AR) decoding strategy (predicting layer 1, then layer 2, then layer 3, then layer 4 for frame 1, before moving to frame 2), what is the total sequence length of tokens the LLM must process for a 5-second audio clip? (a) 250 tokens (b) 1000 tokens (c) 1250 tokens (d) 5000 tokens
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 8 MCQ · 1.0 marks

Imagine that you design a closed-loop system where speech is synthesized from text (TTS) and then transcribed back into text (ASR). Which evaluation measure best captures the accuracy of the text reconstruction at the end of this loop? (a) Signal-to-noise ratio (SNR) of the reconstructed audio (b) Cosine similarity between mel-spectrograms of original and reconstructed audio (c) Perplexity of the reconstructed text under a language model (d) Word Error Rate (WER) between the original and reconstructed text
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 9 MCQ · 1.0 marks

In the CELP algorithm, what does the ”Codebook index” specifically represent? (a) The shape of the vocal tract (b) The index of a pre-defined excitation signal that best matches the residual (c) The volume level of the speaker (d) The duration of the phoneme
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 10 MCQ · 1.0 marks

What is the mathematical advantage of RVQ over standard VQ for a high-dimensional vector space? (a) RVQ eliminates the search time for the closest centroid (b) VQ requires K\(^{N}\) codebook entries for the same precision that RVQ achieves with K × N entries (c) RVQ ensures the reconstructed vector is always mathematically identical to the original (d) RVQ works better with non-linear activation functions in the encoder
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 11 MCQ · 1.0 marks

Which system is most likely to fail if the same sentence is spoken by two different speakers with identical accents? (a) ASR (b) Speaker Verification (c) TTS (d) MT
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 12 MCQ · 1.0 marks

Why is predicting the immediate next frame (e.g., at t+1) often considered a ”trivial” task in speech processing? (a) Speech signals are completely random. (b) Digital speech has very low sample rates. (c) Closely occurring frames (10ms apart) in speech are highly correlated. (d) Next-token prediction is only effective for text, not audio.
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 13 MCQ · 1.0 marks

In RNN-T, if the Joiner network uses a simple addition operation [[IMAGE:3ccef696db85a344_6_2]] , what happens if the Encoder output [[IMAGE:3ccef696db85a344_6_3]] becomes zero? (a) The model becomes a pure CTC model. (b) The model behaves like a pure unconditional Language Model. (c) The model fails to converge. (d) The output becomes random.
Source diagram or notationSource diagram or notation
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 14 MCQ · 1.0 marks

If we convert speech to pure text, what information is primarily lost? (a) The literal words spoken (b) The grammatical structure (c) Speaker identity and emotion (d) The language of the speaker
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 15 MCQ · 1.0 marks

The SUPERB challenge is a benchmark specifically designed to evaluate: (a) Only Automatic Speech Recognition (ASR) models (b) Self-supervised models across multiple speech tasks (c) The physical hardware used for training (d) Text-only LLM performance
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 16 MCQ · 2.0 marks

Contrast the architectural paradigms of Wav2Vec 1.0 and Wav2Vec 2.0. How does the shift from a strictly autoregressive approach (Wav2Vec 1.0) to a bidirectional masked approach (Wav2Vec 2.0) impact the contextual representations learned?
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 17 MCQ · 2.0 marks

When applying K-Means clustering for speech discretization, a specific vocabulary size (number of clusters, K) must be defined. Discuss the trade-offs between choosing a very small vocabulary (e.g., K=50) versus a very large vocabulary (e.g., K=2000). How does this choice impact phonetic discriminability and token sequence bit-rate?
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 18 MCQ · 2.0 marks

Given a latent vector: z=[1.2,0.8] and codebook vectors: c1=[1,1], c2=[0,0] (a) Find the closest codebook vector. (b) Compute the quantized output.
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 19 MCQ · 3.0 marks

Describe the process of modality alignment in a SpeechLLM incase of: (a) SpeechGPT (b) SLAM-ASR Explain with a simple architecture diagram.
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.

Question 20 MCQ · 3.0 marks

Describe the process of how a multi-layer 1D CNN block reduces the sequence length of high- frequency raw audio (e.g., 16 kHz) into a manageable sequence of latent representations. A single 1D CNN layer in a feature extractor has a kernel size of K=10, a stride of S=5, and zero padding P=0. If the input raw audio segment consists of I=400 samples, calculate the temporal dimension, O of the output feature map using the standard convolution formula.
  1. I have answered this question in the offline paper
  2. I have skipped this question
  3. This question does not apply to me

A published solution is not available for this question yet.