da5013_2026T1_Q2_NA.pdf
Deep Learning Practice · Quiz 2 · Jan 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 4.0 marks
[[IMAGE:4bffdb81b4baf18e_2_2]]
The tensor [[IMAGE:4bffdb81b4baf18e_2_3]] has shape:
[[IMAGE:4bffdb81b4baf18e_2_4]]
When using a pre-trained Transformer-based model (such as Wav2Vec 2.0) for a downstream
classification task like Language Identification, what is the primary purpose of applying Mean
Pooling across the sequence dimension?



To reduce the sampling rate of the audio to make it compatible with standard
deep learning layers.
To convert a variable-length sequence of hidden vectors into a single fixed-
length representation that summarizes the entire utterance.
To ensure that the model focuses only on the most intense amplitude peaks of
the speech signal.
To reverse the effects of the Convolutional Neural Network (CNN) encoder and
return the data to the time domain.
A published solution is not available for this question yet.
Question 3 MCQ · 4.0 marks
When using Wav2Vec2Model to extract hidden states for a downstream task, you set
[[IMAGE:4bffdb81b4baf18e_3_5]] in the model call. In PyTorch, what is the most efficient way to
ensure the model does not calculate gradients or update weights during this feature extraction
process, thereby saving memory and computation?

Use [[IMAGE:4bffdb81b4baf18e_3_6]] before passing the audio through the model.

Wrap the extraction code block with the [[IMAGE:4bffdb81b4baf18e_3_7]] : context manager.

Manually set each layer's [[IMAGE:4bffdb81b4baf18e_3_8]] attribute to False using a for loop
before every forward pass.

Use the [[IMAGE:4bffdb81b4baf18e_3_9]] function on the output tensor to remove
unnecessary dimensions.

A published solution is not available for this question yet.
Question 4 MCQ · 4.0 marks
When fine-tuning a Whisper model using the Hugging Face transformers library, which of the
following statements best describes the role and behavior of the WhisperProcessor?
It is a standalone model that converts audio directly into text before passing it
to the WhisperForConditionalGeneration class.
It is a wrapper class that combines a WhisperFeatureExtractor (for audio) and
a WhisperTokenizer (for text) into a single object to simplify data preprocessing and padding.
It is primarily used to change the sampling_rate of the raw audio files to
16,000 Hz during the dataset.map() phase.
It is a specialized loss function that calculates the Word Error Rate (WER) by
comparing the model's audio features directly against the ground-truth text label
A published solution is not available for this question yet.
Question 5 MCQ · 4.0 marks
When working with TTS models like SpeechT5, what is the role of the speaker_embeddings?
To translate the text into different languages before speaking.
To define the specific vocal characteristics (voice identity) of the speaker.
To remove background noise from the generated file.
To increase the speed of the audio generation.
A published solution is not available for this question yet.
Question 6 NAT · 3.0 marks
Consider the following code snippet:
[[IMAGE:4bffdb81b4baf18e_4_10]]
Assume that the audio sample stored in [[IMAGE:4bffdb81b4baf18e_4_11]] has:
• [[IMAGE:4bffdb81b4baf18e_4_12]] : 48,000
• [[IMAGE:4bffdb81b4baf18e_4_13]] : 13,440,000
What is the duration of the audio in seconds?




A published solution is not available for this question yet.
Question 7 NAT · 3.0 marks
Calculate the Word Error Rate (WER) for the following example:
Actual Sentence : [[IMAGE:4bffdb81b4baf18e_5_14]]
ASR prediction : [[IMAGE:4bffdb81b4baf18e_5_15]]


A published solution is not available for this question yet.
Question 8 MSQ · 4.0 marks
In a standard modular speaker diarization pipeline, which of the following processes are
responsible for identifying and grouping unique speaker characteristics from speech segments?
Speaker Embedding Extraction: Generating numerical representations (e.g., x-
vectors) that capture the identity-bearing features of the vocal signal.
Part-of-Speech (POS) Tagging: Assigning grammatical categories to words to
determine the syntactic structure of the conversation.
Clustering: Applying algorithms like Spectral Clustering or AHC to group
embeddings into clusters corresponding to individual speakers.
Named Entity Recognition (NER): Extracting proper nouns and specific entities
from the transcribed text to identify participants by name.
Distance/Similarity Metric Computation: Calculating the mathematical
relationship (e.g., Cosine similarity or distances) between different speech segments.
A published solution is not available for this question yet.
Question 9 MSQ · 4.0 marks
Whisper can be used as a component in a diarization pipeline in which of the following ways?
As the automatic speech recognition backend for transcription after
diarization
To provide word-level timestamps for aligning speaker segments
As the primary clustering algorithm for speaker embeddings
As the speaker embedding extractor
A published solution is not available for this question yet.
Question 10 MSQ · 4.0 marks
When implementing diarization with SpeechBrain and PyTorch, which of the following practices
are recommended? (Select all that apply)
Use .to(device) to move models to GPU
Normalize audio waveforms before feeding to the embedding model
Use batch processing for extracting embeddings from multiple segments
Pass raw text tokens to the speaker encoder
Use torch.no_grad() during inference to save memory
A published solution is not available for this question yet.
Question 11 MCQ · 2.0 marks
You are building an ASR pipeline for Assamese. Complete the missing parts of the script to ensure
the audio is resampled, features are extracted, and the trainer is initialized correctly.
[[IMAGE:4bffdb81b4baf18e_7_16]]
Based on the above data, answer the given subquestions.
Choose the option from below to fill in the blank at A.

apply
cast_column
transform
batch_decode
A published solution is not available for this question yet.
Question 12 MCQ · 2.0 marks
You are building an ASR pipeline for Assamese. Complete the missing parts of the script to ensure
the audio is resampled, features are extracted, and the trainer is initialized correctly.
[[IMAGE:4bffdb81b4baf18e_7_16]]
Based on the above data, answer the given subquestions.
Choose the option from below to fill in the blank at B.

WhisperTokenizer
WhisperProcessor
WhisperForConditionalGeneration
Seq2SeqTrainingArguments
A published solution is not available for this question yet.
Question 13 MCQ · 2.0 marks
You are building an ASR pipeline for Assamese. Complete the missing parts of the script to ensure
the audio is resampled, features are extracted, and the trainer is initialized correctly.
[[IMAGE:4bffdb81b4baf18e_7_16]]
Based on the above data, answer the given subquestions.
Choose the option from below to fill in the blank at C.

feature_extractor
tokenizer
collator
None of these
A published solution is not available for this question yet.
Question 14 MCQ · 3.0 marks
[[IMAGE:4bffdb81b4baf18e_9_17]]
**output of print(vocabs) is as given below**
[[IMAGE:4bffdb81b4baf18e_9_18]]
Based on the above data, answer the given subquestions.
For the given code what needs to be filled in place of **missing code** to get the vocabulary which
consists of only unique character ?


[[IMAGE:4bffdb81b4baf18e_10_19]]

[[IMAGE:4bffdb81b4baf18e_10_20]]

[[IMAGE:4bffdb81b4baf18e_10_21]]

[[IMAGE:4bffdb81b4baf18e_10_22]]

A published solution is not available for this question yet.
Question 15 NAT · 2.0 marks
[[IMAGE:4bffdb81b4baf18e_9_17]]
**output of print(vocabs) is as given below**
[[IMAGE:4bffdb81b4baf18e_9_18]]
Based on the above data, answer the given subquestions.
If [[IMAGE:4bffdb81b4baf18e_10_23]] is generated correctly what would be the length of it ?



A published solution is not available for this question yet.
Question 16 MCQ · 3.0 marks
In the context of Hugging Face TTS, what is the purpose of the 'Vocoder' component?
To translate the input text into a different language.
To tokenize the text into sub-word units.
To convert intermediate acoustic features (like mel-spectrograms) into audible
waveforms.
To compress the final audio file into a ZIP format.
A published solution is not available for this question yet.
Question 17 MCQ · 2.0 marks
Why is it important to normalize or resample reference audio to 16,000 Hz when extracting
speaker embeddings for models like SpeechT5?
To make the file size as large as possible.
Because Python cannot process any frequency higher than 16kHz.
Because 16kHz is the highest quality audio humans can hear.
Because the model was specifically trained on 16kHz audio data.
A published solution is not available for this question yet.