da5013_2026T2_Q2_NA.pdf
Deep Learning Practice · Quiz 2 · May 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MSQ · 3.0 marks
Which of the following factors can adversely impact the accuracy of speech language
identification?
Background noise
Very short utterances
Code-switching
Large balanced training data
A published solution is not available for this question yet.
Question 3 MSQ · 3.0 marks
You are building an end-to-end speaker-attributed transcription pipeline using the Whisper model
for speech recognition and timestamp generation, the ECAPA-TDNN model for speaker
embedding extraction, and the SpeechBrain toolkit to integrate the diarization pipeline.
Which of the following subtasks are necessary to achieve this? (Select all that apply)
Voice Activity Detection (VAD) / Segmentation
Speech Recognition and Timestamp Generation
Speaker Embedding Extraction
Language Identification
Speaker Clustering
Text-to-Speech (TTS) Synthesis
Alignment and Segment Merging
A published solution is not available for this question yet.
Question 4 MSQ · 3.0 marks
When feeding audio into the pretrained speechbrain/spkrec-ecapa-voxceleb model, which of the
following statements regarding the expected input is (are) true? (Select all that apply)
The system is trained on single-channel (mono) recordings.
The user must manually extract acoustic features (like MFCCs or Mel-
filterbanks) before passing the data to the model.
The expected sampling rate for the input audio is 16kHz.
The model accepts variable-length audio inputs, dynamically generating fixed-
size embeddings regardless of the utterance duration.
A published solution is not available for this question yet.
Question 5 MSQ · 3.0 marks
Which statements accurately describe the roles, constraints, and data-processing behaviors when
preparing batches for speech models (e.g., Whisper, Wav2Vec2, and SpeechT5) using the Hugging
Face Transformers library?
For sequence-to-sequence ASR models such as Whisper, the data collator
replaces padded token label IDs with [[IMAGE:60e13f4012e51d90_3_2]] so they are ignored by the cross-entropy loss during
training.

Dynamic padding in speech data collators pads input features and target
labels only to the maximum sequence length within each batch, reducing unnecessary memory
usage compared to dataset-wide padding.
Speech data collators automatically resample raw audio (e.g., from 44.1 kHz to
16 kHz) during batch collation before creating model inputs.
During SpeechT5 TTS training, the target spectrogram is padded to a length
that is a multiple of the reduction factor ( [[IMAGE:60e13f4012e51d90_3_3]] ) so that it is compatible with the decoder's frame-
reduction architecture.

[[IMAGE:60e13f4012e51d90_3_4]] computes log-Mel spectrogram features
directly from raw audio waveforms while constructing each batch.

A published solution is not available for this question yet.
Question 6 NAT · 2.0 marks
An audio recording is sampled at 16,000 Hz for 8 seconds. How many samples are contained in the
recording?
A published solution is not available for this question yet.
Question 7 NAT · 2.0 marks
A multilingual language identification model is trained to classify **40 languages**. The training
dataset contains **1000 audio clips per language**, and each audio clip has been resampled to **16**
**kHz**. The Transformer encoder produces a **1024 dimensional embedding** for every audio clip.
During training, a mini-batch contains **64 audio clips**. Before the classification layer, the
embeddings are stacked into a tensor of shape **(batch_size, embedding_dimension)**.
**Question:** Determine the **total number of embedding values** in this output tensor.
A published solution is not available for this question yet.
Question 8 NAT · 2.0 marks
An Automatic Speech Recognition (ASR) system produces a hypothesis transcript for a spoken
audio clip. Calculate the Word Error Rate (WER) of the system given the reference (ground truth)
and hypothesis transcripts below:
• Reference (Ground Truth): "the quick brown fox jumps over the lazy black dog"
• Hypothesis (ASR Output): "the fast brown fox jumped over lazy black dog today"
(Note: Provide your answer as a decimal rounded to two decimal places.)
A published solution is not available for this question yet.
Question 9 MCQ · 2.0 marks
In a spoken language identification task, a pretrained speech model is used to extract features
from an input audio waveform. To obtain a single fixed-length embedding representing the entire
audio clip for language classification, fill in the missing statement in the following code.
[[IMAGE:60e13f4012e51d90_5_5]]

sum
max
mean
flatten
A published solution is not available for this question yet.
Question 10 MCQ · 2.0 marks
What is the effect of setting compute_type="int8" in WhisperModel(whisper_model,
compute_type="int8")?
It restricts the model to only transcribing 8-second audio chunks.
It forces the audio file to be read as an 8-bit WAV file.
It quantizes the model weights to 8-bit integers, significantly reducing
GPU/CPU memory usage with minimal accuracy loss.
It limits the language detection capability to the 8 most common languages.
A published solution is not available for this question yet.
Question 11 MCQ · 2.0 marks
Consider the following incomplete script intended for fine-tuning a Wav2Vec2 model:
[[IMAGE:60e13f4012e51d90_6_6]]
Which of the following code blocks correctly fills in the blank to ensure the model's output layer
matches the tokenizer's vocabulary size and recognizes its padding token?

pad_token_id=tokenizer.pad_token_id, vocab_size=len(tokenizer)
padding_id=tokenizer.pad_token, vocabulary=tokenizer.vocab
pad_token=tokenizer.pad_token_id, size=tokenizer.vocab_size
pad_token_id=tokenizer.pad_token, vocab_len=len(tokenizer)
A published solution is not available for this question yet.
Question 12 MCQ · 2.0 marks
Which of the following describes the correct input for the Decoder during the training of a
Seq2Seq ASR model like Whisper?
The raw audio waveform normalized to a zero mean.
A sequence of random noise to improve robustness.
The Log-Mel Spectrogram extracted from the audio signal.
The tokens predicted by the decoder up to the previous time step during
inference
A published solution is not available for this question yet.
Question 13 MCQ · 2.0 marks
When fine-tuning a pre-trained Transformer-based ASR model using Connectionist Temporal
Classification (CTC), what is the primary function of the 'blank' token?
Indicating the end of a sentence.
Handling out-of-vocabulary words.
Aligning variable-length sequences.
Representing silent audio segments.
A published solution is not available for this question yet.
Question 14 MCQ · 2.0 marks
Consider the following PyTorch implementation of a **Convolutional Neural Network (CNN)**
designed for an **audio language identification** task. The input to the model consists of speech
embeddings with shape [[IMAGE:60e13f4012e51d90_7_7]] .
[[IMAGE:60e13f4012e51d90_8_8]]
Based on the above data, answer the given subquestions.
Choose the option to fill in the blanks for the missing code (A)


[[IMAGE:60e13f4012e51d90_8_9]]

[[IMAGE:60e13f4012e51d90_8_10]]

[[IMAGE:60e13f4012e51d90_9_11]]

[[IMAGE:60e13f4012e51d90_9_12]]

[[IMAGE:60e13f4012e51d90_9_13]]

A published solution is not available for this question yet.
Question 15 NAT · 3.0 marks
Consider the following PyTorch implementation of a **Convolutional Neural Network (CNN)**
designed for an **audio language identification** task. The input to the model consists of speech
embeddings with shape [[IMAGE:60e13f4012e51d90_7_7]] .
[[IMAGE:60e13f4012e51d90_8_8]]
Based on the above data, answer the given subquestions.
The input tensor **train_data** has shape
[[IMAGE:60e13f4012e51d90_9_14]]
Determine the correct value of **in_features** for the **fc1** layer.



A published solution is not available for this question yet.
Question 16 MCQ · 3.0 marks
[[IMAGE:60e13f4012e51d90_9_15]] returns a tensor of shape [[IMAGE:60e13f4012e51d90_9_16]] . For a stereo file (2
channels), the tensor shape is [[IMAGE:60e13f4012e51d90_9_17]] .
If you pass this tensor directly into [[IMAGE:60e13f4012e51d90_9_18]] without first averaging the channels down
to mono, how will the model interpret it?
[[IMAGE:60e13f4012e51d90_10_19]]





It will raise a [[IMAGE:60e13f4012e51d90_10_20]] , since the ECAPA-TDNN model strictly requires
mono (1-channel) input.

It will automatically mean-pool the left and right channels into a mono signal
before processing.
It will treat the tensor's first dimension as a batch dimension, producing two
separate embeddings — one per channel.
It will produce a single embedding, but with double the feature dimension to
account for the extra channel.
A published solution is not available for this question yet.
Question 17 MCQ · 3.0 marks
Why might you freeze the encoder layers of a pre-trained ASR model during the initial phases of
fine-tuning on a very small dataset ?
Reducing the inference latency.
Eliminating the need for CTC.
Preventing catastrophic forgetting.
Increasing the model capacity.
A published solution is not available for this question yet.
Question 18 MCQ · 3.0 marks
What is the primary disadvantage of using a character-level tokenizer compared to a Byte-Pair
Encoding (BPE) subword tokenizer for ASR fine-tuning ?
Longer output sequence lengths.
Larger vocabulary size.
Inability to predict new words.
Higher memory usage for embedding.
A published solution is not available for this question yet.
Question 19 MCQ · 3.0 marks
In a standard TTS pipeline involving an acoustic model and a vocoder, which component is
primarily responsible for converting Mel-spectrograms into time-domain waveforms during the
fine-tuning process?
The Vocoder.
The Duration Predictor.
The Phonemizer.
The Attention Mechanism.
A published solution is not available for this question yet.
Question 20 MSQ · 2.0 marks
Which of the following linkages is possible in agglomerative clustering?
Ward Linkage
Web Linkage
Single Linkage
Hierarchical Linkage
A published solution is not available for this question yet.
Question 21 NAT · 3.0 marks
The "sentence-transformers/all-mpnet-base-v2" model generates **768-dimensional** sentence
embeddings. A data processing pipeline processes **400 unique documents, each containing 25**
**sentences**. It stores the embedding vector for every sentence in a single contiguous NumPy array.
How many floating-point values does the array contain?
A published solution is not available for this question yet.