MauryaHub PYQ Practice

da5013_2026T1_Q2_NA.pdf

Deep Learning Practice · Quiz 2 · Jan 2026

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 4.0 marks

[[IMAGE:4bffdb81b4baf18e_2_2]] The tensor [[IMAGE:4bffdb81b4baf18e_2_3]] has shape: [[IMAGE:4bffdb81b4baf18e_2_4]] When using a pre-trained Transformer-based model (such as Wav2Vec 2.0) for a downstream classification task like Language Identification, what is the primary purpose of applying Mean Pooling across the sequence dimension?
Source diagram or notationSource diagram or notationSource diagram or notation
  1. To reduce the sampling rate of the audio to make it compatible with standard deep learning layers.
  2. To convert a variable-length sequence of hidden vectors into a single fixed- length representation that summarizes the entire utterance.
  3. To ensure that the model focuses only on the most intense amplitude peaks of the speech signal.
  4. To reverse the effects of the Convolutional Neural Network (CNN) encoder and return the data to the time domain.

A published solution is not available for this question yet.

Question 3 MCQ · 4.0 marks

When using Wav2Vec2Model to extract hidden states for a downstream task, you set [[IMAGE:4bffdb81b4baf18e_3_5]] in the model call. In PyTorch, what is the most efficient way to ensure the model does not calculate gradients or update weights during this feature extraction process, thereby saving memory and computation?
Source diagram or notation
  1. Use [[IMAGE:4bffdb81b4baf18e_3_6]] before passing the audio through the model.
    Source diagram or notation
  2. Wrap the extraction code block with the [[IMAGE:4bffdb81b4baf18e_3_7]] : context manager.
    Source diagram or notation
  3. Manually set each layer's [[IMAGE:4bffdb81b4baf18e_3_8]] attribute to False using a for loop before every forward pass.
    Source diagram or notation
  4. Use the [[IMAGE:4bffdb81b4baf18e_3_9]] function on the output tensor to remove unnecessary dimensions.
    Source diagram or notation

A published solution is not available for this question yet.

Question 4 MCQ · 4.0 marks

When fine-tuning a Whisper model using the Hugging Face transformers library, which of the following statements best describes the role and behavior of the WhisperProcessor?
  1. It is a standalone model that converts audio directly into text before passing it to the WhisperForConditionalGeneration class.
  2. It is a wrapper class that combines a WhisperFeatureExtractor (for audio) and a WhisperTokenizer (for text) into a single object to simplify data preprocessing and padding.
  3. It is primarily used to change the sampling_rate of the raw audio files to 16,000 Hz during the dataset.map() phase.
  4. It is a specialized loss function that calculates the Word Error Rate (WER) by comparing the model's audio features directly against the ground-truth text label

A published solution is not available for this question yet.

Question 5 MCQ · 4.0 marks

When working with TTS models like SpeechT5, what is the role of the speaker_embeddings?
  1. To translate the text into different languages before speaking.
  2. To define the specific vocal characteristics (voice identity) of the speaker.
  3. To remove background noise from the generated file.
  4. To increase the speed of the audio generation.

A published solution is not available for this question yet.

Question 6 NAT · 3.0 marks

Consider the following code snippet: [[IMAGE:4bffdb81b4baf18e_4_10]] Assume that the audio sample stored in [[IMAGE:4bffdb81b4baf18e_4_11]] has: • [[IMAGE:4bffdb81b4baf18e_4_12]] : 48,000 • [[IMAGE:4bffdb81b4baf18e_4_13]] : 13,440,000 What is the duration of the audio in seconds?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 7 NAT · 3.0 marks

    Calculate the Word Error Rate (WER) for the following example: Actual Sentence : [[IMAGE:4bffdb81b4baf18e_5_14]] ASR prediction : [[IMAGE:4bffdb81b4baf18e_5_15]]
    Source diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 8 MSQ · 4.0 marks

      In a standard modular speaker diarization pipeline, which of the following processes are responsible for identifying and grouping unique speaker characteristics from speech segments?
      1. Speaker Embedding Extraction: Generating numerical representations (e.g., x- vectors) that capture the identity-bearing features of the vocal signal.
      2. Part-of-Speech (POS) Tagging: Assigning grammatical categories to words to determine the syntactic structure of the conversation.
      3. Clustering: Applying algorithms like Spectral Clustering or AHC to group embeddings into clusters corresponding to individual speakers.
      4. Named Entity Recognition (NER): Extracting proper nouns and specific entities from the transcribed text to identify participants by name.
      5. Distance/Similarity Metric Computation: Calculating the mathematical relationship (e.g., Cosine similarity or distances) between different speech segments.

      A published solution is not available for this question yet.

      Question 9 MSQ · 4.0 marks

      Whisper can be used as a component in a diarization pipeline in which of the following ways?
      1. As the automatic speech recognition backend for transcription after diarization
      2. To provide word-level timestamps for aligning speaker segments
      3. As the primary clustering algorithm for speaker embeddings
      4. As the speaker embedding extractor

      A published solution is not available for this question yet.

      Question 10 MSQ · 4.0 marks

      When implementing diarization with SpeechBrain and PyTorch, which of the following practices are recommended? (Select all that apply)
      1. Use .to(device) to move models to GPU
      2. Normalize audio waveforms before feeding to the embedding model
      3. Use batch processing for extracting embeddings from multiple segments
      4. Pass raw text tokens to the speaker encoder
      5. Use torch.no_grad() during inference to save memory

      A published solution is not available for this question yet.

      Question 11 MCQ · 2.0 marks

      You are building an ASR pipeline for Assamese. Complete the missing parts of the script to ensure the audio is resampled, features are extracted, and the trainer is initialized correctly. [[IMAGE:4bffdb81b4baf18e_7_16]] Based on the above data, answer the given subquestions.
      Choose the option from below to fill in the blank at A.
      Source diagram or notation
      1. apply
      2. cast_column
      3. transform
      4. batch_decode

      A published solution is not available for this question yet.

      Question 12 MCQ · 2.0 marks

      You are building an ASR pipeline for Assamese. Complete the missing parts of the script to ensure the audio is resampled, features are extracted, and the trainer is initialized correctly. [[IMAGE:4bffdb81b4baf18e_7_16]] Based on the above data, answer the given subquestions.
      Choose the option from below to fill in the blank at B.
      Source diagram or notation
      1. WhisperTokenizer
      2. WhisperProcessor
      3. WhisperForConditionalGeneration
      4. Seq2SeqTrainingArguments

      A published solution is not available for this question yet.

      Question 13 MCQ · 2.0 marks

      You are building an ASR pipeline for Assamese. Complete the missing parts of the script to ensure the audio is resampled, features are extracted, and the trainer is initialized correctly. [[IMAGE:4bffdb81b4baf18e_7_16]] Based on the above data, answer the given subquestions.
      Choose the option from below to fill in the blank at C.
      Source diagram or notation
      1. feature_extractor
      2. tokenizer
      3. collator
      4. None of these

      A published solution is not available for this question yet.

      Question 14 MCQ · 3.0 marks

      [[IMAGE:4bffdb81b4baf18e_9_17]] **output of print(vocabs) is as given below** [[IMAGE:4bffdb81b4baf18e_9_18]] Based on the above data, answer the given subquestions.
      For the given code what needs to be filled in place of **missing code** to get the vocabulary which consists of only unique character ?
      Source diagram or notationSource diagram or notation
      1. [[IMAGE:4bffdb81b4baf18e_10_19]]
        Source diagram or notation
      2. [[IMAGE:4bffdb81b4baf18e_10_20]]
        Source diagram or notation
      3. [[IMAGE:4bffdb81b4baf18e_10_21]]
        Source diagram or notation
      4. [[IMAGE:4bffdb81b4baf18e_10_22]]
        Source diagram or notation

      A published solution is not available for this question yet.

      Question 15 NAT · 2.0 marks

      [[IMAGE:4bffdb81b4baf18e_9_17]] **output of print(vocabs) is as given below** [[IMAGE:4bffdb81b4baf18e_9_18]] Based on the above data, answer the given subquestions.
      If [[IMAGE:4bffdb81b4baf18e_10_23]] is generated correctly what would be the length of it ?
      Source diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 16 MCQ · 3.0 marks

        In the context of Hugging Face TTS, what is the purpose of the 'Vocoder' component?
        1. To translate the input text into a different language.
        2. To tokenize the text into sub-word units.
        3. To convert intermediate acoustic features (like mel-spectrograms) into audible waveforms.
        4. To compress the final audio file into a ZIP format.

        A published solution is not available for this question yet.

        Question 17 MCQ · 2.0 marks

        Why is it important to normalize or resample reference audio to 16,000 Hz when extracting speaker embeddings for models like SpeechT5?
        1. To make the file size as large as possible.
        2. Because Python cannot process any frequency higher than 16kHz.
        3. Because 16kHz is the highest quality audio humans can hear.
        4. Because the model was specifically trained on 16kHz audio data.

        A published solution is not available for this question yet.