da5013_2025T3_Q2_NA.pdf
Deep Learning Practice · Quiz 2 · Sep 2025
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 197 NAT · 5.0 marks
Using every-visit Monte Carlo, estimate [[IMAGE:0b0841408a2ed374_1_0]] and [[IMAGE:0b0841408a2ed374_1_1]] . What is the value of [[IMAGE:0b0841408a2ed374_1_2]] ? (use
[[IMAGE:0b0841408a2ed374_1_3]] )




A published solution is not available for this question yet.
Question 198 MCQ · 2.0 marks
Batch Monte Carlo (MC) methods estimate the value function by finding the least-squares fit to the
sampled returns generated under the policy.
True
False
**DLP**
**Section Id :** 640653121914
**Section Number :** 9
**Section type :** Online
**Mandatory or Optional :** Mandatory
**Number of Questions :** 16
**Number of Questions to be attempted :** 16
**Section Marks :** 50
**Display Number Panel :** Yes
**Section Negative Marks :** 0
**Group All Questions :** No
**Enable Mark as Answered Mark for Review and**
No
**Clear Response :**
**Section Maximum Duration :** 0
**Section Minimum Duration :** 0
**Section Time In :** Minutes
**Maximum Instruction Time :** 0
A published solution is not available for this question yet.
Question 200 MCQ · 3.0 marks
In an embedding-based speaker diarization pipeline, what is the primary role of the
Agglomerative Clustering algorithm?
To transcribe the audio segments into text using a model like Whisper.
To extract a 512-dimensional embedding (x-vector) from each audio segment.
To compare the cosine distance between segment embeddings and iteratively
merge the closest ones until a target number of speakers is reached.
To detect non-speech segments (Voice Activity Detection) and discard them.
A published solution is not available for this question yet.
Question 201 MCQ · 3.0 marks
In modern Text-to-Speech (TTS) pipelines like Tacotron2 or FastSpeech, what is the specific role of
a component like HiFi-GAN or WaveNet?
To convert the input text into a sequence of phonemes (G2P).
To predict the duration of each phoneme in the sequence.
To convert the intermediate mel-spectrogram representation into a raw audio
waveform.
To extract a speaker embedding from a reference audio file.
A published solution is not available for this question yet.
Question 202 MCQ · 3.0 marks
In the Wav2Vec2 pipeline, a 1-second (16000 samples) audio clip is passed to the 'Wav2Vec2Model'
and results in a 'last hidden state' of shape (1, 49, 768). What do the dimensions 49 and 768
represent?
49 = number of possible phonemes; 768 = batch size.
49 = the model’s hidden dimension; 768 = the downsampled sequence length.
49 = the downsampled sequence length (timesteps); 768 = the model’s hidden
feature dimension.
49 = number of attention heads; 768 = the vocabulary size.
A published solution is not available for this question yet.
Question 203 MCQ · 3.0 marks
In an ASR model like Wav2Vec2-CTC, the final linear layer outputs a tensor of logits. What does this
tensor represent?
The final transcribed text string.
A probability distribution over potential speaker identities.
Raw, unnormalized scores for each token in the vocabulary (including the
blank token) for each time step.
The mel-spectrogram of the input audio, compressed by the encoder.
A published solution is not available for this question yet.
Question 204 MCQ · 3.0 marks
What is the primary architectural innovation of FastSpeech that makes it faster and more robust
than an auto-regressive model like Tacotron2?
It uses a more powerful vocoder (HiFi-GAN) to generate speech.
It predicts raw audio directly instead of mel-spectrograms.
It replaces the auto-regressive attention mechanism with a parallel ”Duration
Predictor” and ”Length Regulator”.
It uses a much larger Transformer encoder to understand text.
A published solution is not available for this question yet.
Question 205 MCQ · 3.0 marks
An audio clip has a duration of 5 seconds and is recorded at a sample rate of 44.1 kHz. What
happens to the total number of samples if the audio is downsampled to 22.05 kHz?
The number of samples remains the same.
The number of samples is halved.
The number of samples doubles.
The number of samples becomes one-quarter of the original.
A published solution is not available for this question yet.
Question 206 MCQ · 2.0 marks
[[IMAGE:0b0841408a2ed374_4_4]]

[[IMAGE:0b0841408a2ed374_4_5]]

[[IMAGE:0b0841408a2ed374_4_6]]

[[IMAGE:0b0841408a2ed374_4_7]]

[[IMAGE:0b0841408a2ed374_4_8]]

A published solution is not available for this question yet.
Question 207 MSQ · 4.0 marks
A complete speaker diarization pipeline is used to determine "who spoke when". Which of the
following are essential components of a modern, embedding-based diarization system? (Select
ALL that apply)
Voice Activity Detection (VAD) to filter out silence.
A speaker embedding model (e.g., ECAPA-TDNN) to create vectors for speech
segments.
A clustering algorithm (e.g., Agglomerative Clustering) to group segments by
speaker.
A Text-to-Speech (TTS) engine to generate the final output.
A published solution is not available for this question yet.
Question 208 MSQ · 4.0 marks
[[IMAGE:0b0841408a2ed374_5_9]]

[[IMAGE:0b0841408a2ed374_5_10]]

[[IMAGE:0b0841408a2ed374_5_11]]

[[IMAGE:0b0841408a2ed374_5_12]]

[[IMAGE:0b0841408a2ed374_5_13]]

A published solution is not available for this question yet.
Question 209 MSQ · 4.0 marks
Which of the following statements accurately describe the Wav2Vec2 model and its components?
(Select ALL that apply)
[[IMAGE:0b0841408a2ed374_5_14]]

[[IMAGE:0b0841408a2ed374_5_15]]

[[IMAGE:0b0841408a2ed374_6_16]]

[[IMAGE:0b0841408a2ed374_6_17]]

A published solution is not available for this question yet.
Question 210 MSQ · 4.0 marks
Regarding the SpeechT5 model for Text-to-Speech (TTS), which of the following statements are
true? (Select ALL that apply)
It is an encoder-decoder Transformer model.
To generate speech for a specific voice, it requires a speaker embedding as an
additional input.
HiFi-GAN is an integrated and mandatory part of the SpeechT5 model
architecture.
It is pre-trained only on text data, similar to BERT.
A published solution is not available for this question yet.
Question 211 MSQ · 4.0 marks
[[IMAGE:0b0841408a2ed374_6_18]]

[[IMAGE:0b0841408a2ed374_6_19]]

[[IMAGE:0b0841408a2ed374_6_20]]

[[IMAGE:0b0841408a2ed374_6_21]]

[[IMAGE:0b0841408a2ed374_6_22]]

A published solution is not available for this question yet.
Question 212 NAT · 3.0 marks
You are evaluating an ASR model using 'jiwer'. The ground truth transcription is: "the quick brown
fox jumps over the lazy dog" (9 words). The model predicts: "the quick fox jumped over a lazy
dog". Calculate the Word Error Rate (WER). Provide the answer as a decimal rounded to two
places.
A published solution is not available for this question yet.
Question 213 NAT · 3.0 marks
A 3.5-second audio file is sampled at 16000 Hz. It is fed into a 'Wav2Vec2Model'
("facebook/wav2vec2-base-960h"), which uses a CNN feature extractor that downsamples the
input by a factor of 320. What will be the sequence length (i.e., the number of time steps) of the
'last hidden state' output?
A published solution is not available for this question yet.
Question 214 NAT · 4.0 marks
A speaker diarization script processes a dataset with 200 unique speakers. It extracts an average
of 15 segments for each speaker using a model, which produces 512- dimensional embeddings. If
each embedding vector is stored using 32-bit floating-point precision (4 bytes per value), what is
the total memory required to store all the segment embeddings in Megabytes (MB)? (Assume 1
MB = 1024\(^{2}\) bytes). Round your answer to the nearest integer.
A published solution is not available for this question yet.