MauryaHub PYQ Practice

da5013_2025T3_Q2_NA.pdf

Deep Learning Practice · Quiz 2 · Sep 2025

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 197 NAT · 5.0 marks

Using every-visit Monte Carlo, estimate [[IMAGE:0b0841408a2ed374_1_0]] and [[IMAGE:0b0841408a2ed374_1_1]] . What is the value of [[IMAGE:0b0841408a2ed374_1_2]] ? (use [[IMAGE:0b0841408a2ed374_1_3]] )
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 198 MCQ · 2.0 marks

    Batch Monte Carlo (MC) methods estimate the value function by finding the least-squares fit to the sampled returns generated under the policy.
    1. True
    2. False       **DLP** **Section Id :** 640653121914 **Section Number :** 9 **Section type :** Online **Mandatory or Optional :** Mandatory **Number of Questions :** 16 **Number of Questions to be attempted :** 16 **Section Marks :** 50 **Display Number Panel :** Yes **Section Negative Marks :** 0 **Group All Questions :** No **Enable Mark as Answered Mark for Review and** No **Clear Response :** **Section Maximum Duration :** 0 **Section Minimum Duration :** 0 **Section Time In :** Minutes **Maximum Instruction Time :** 0

    A published solution is not available for this question yet.

    Question 200 MCQ · 3.0 marks

    In an embedding-based speaker diarization pipeline, what is the primary role of the Agglomerative Clustering algorithm?
    1. To transcribe the audio segments into text using a model like Whisper.
    2. To extract a 512-dimensional embedding (x-vector) from each audio segment.
    3. To compare the cosine distance between segment embeddings and iteratively merge the closest ones until a target number of speakers is reached.
    4. To detect non-speech segments (Voice Activity Detection) and discard them.

    A published solution is not available for this question yet.

    Question 201 MCQ · 3.0 marks

    In modern Text-to-Speech (TTS) pipelines like Tacotron2 or FastSpeech, what is the specific role of a component like HiFi-GAN or WaveNet?
    1. To convert the input text into a sequence of phonemes (G2P).
    2. To predict the duration of each phoneme in the sequence.
    3. To convert the intermediate mel-spectrogram representation into a raw audio waveform.
    4. To extract a speaker embedding from a reference audio file.

    A published solution is not available for this question yet.

    Question 202 MCQ · 3.0 marks

    In the Wav2Vec2 pipeline, a 1-second (16000 samples) audio clip is passed to the 'Wav2Vec2Model' and results in a 'last hidden state' of shape (1, 49, 768). What do the dimensions 49 and 768 represent?
    1. 49 = number of possible phonemes; 768 = batch size.
    2. 49 = the model’s hidden dimension; 768 = the downsampled sequence length.
    3. 49 = the downsampled sequence length (timesteps); 768 = the model’s hidden feature dimension.
    4. 49 = number of attention heads; 768 = the vocabulary size.

    A published solution is not available for this question yet.

    Question 203 MCQ · 3.0 marks

    In an ASR model like Wav2Vec2-CTC, the final linear layer outputs a tensor of logits. What does this tensor represent?
    1. The final transcribed text string.
    2. A probability distribution over potential speaker identities.
    3. Raw, unnormalized scores for each token in the vocabulary (including the blank token) for each time step.
    4. The mel-spectrogram of the input audio, compressed by the encoder.

    A published solution is not available for this question yet.

    Question 204 MCQ · 3.0 marks

    What is the primary architectural innovation of FastSpeech that makes it faster and more robust than an auto-regressive model like Tacotron2?
    1. It uses a more powerful vocoder (HiFi-GAN) to generate speech.
    2. It predicts raw audio directly instead of mel-spectrograms.
    3. It replaces the auto-regressive attention mechanism with a parallel ”Duration Predictor” and ”Length Regulator”.
    4. It uses a much larger Transformer encoder to understand text.

    A published solution is not available for this question yet.

    Question 205 MCQ · 3.0 marks

    An audio clip has a duration of 5 seconds and is recorded at a sample rate of 44.1 kHz. What happens to the total number of samples if the audio is downsampled to 22.05 kHz?
    1. The number of samples remains the same.
    2. The number of samples is halved.
    3. The number of samples doubles.
    4. The number of samples becomes one-quarter of the original.

    A published solution is not available for this question yet.

    Question 206 MCQ · 2.0 marks

    [[IMAGE:0b0841408a2ed374_4_4]]
    Source diagram or notation
    1. [[IMAGE:0b0841408a2ed374_4_5]]
      Source diagram or notation
    2. [[IMAGE:0b0841408a2ed374_4_6]]
      Source diagram or notation
    3. [[IMAGE:0b0841408a2ed374_4_7]]
      Source diagram or notation
    4. [[IMAGE:0b0841408a2ed374_4_8]]
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 207 MSQ · 4.0 marks

    A complete speaker diarization pipeline is used to determine "who spoke when". Which of the following are essential components of a modern, embedding-based diarization system? (Select ALL that apply)
    1. Voice Activity Detection (VAD) to filter out silence.
    2. A speaker embedding model (e.g., ECAPA-TDNN) to create vectors for speech segments.
    3. A clustering algorithm (e.g., Agglomerative Clustering) to group segments by speaker.
    4. A Text-to-Speech (TTS) engine to generate the final output.

    A published solution is not available for this question yet.

    Question 208 MSQ · 4.0 marks

    [[IMAGE:0b0841408a2ed374_5_9]]
    Source diagram or notation
    1. [[IMAGE:0b0841408a2ed374_5_10]]
      Source diagram or notation
    2. [[IMAGE:0b0841408a2ed374_5_11]]
      Source diagram or notation
    3. [[IMAGE:0b0841408a2ed374_5_12]]
      Source diagram or notation
    4. [[IMAGE:0b0841408a2ed374_5_13]]
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 209 MSQ · 4.0 marks

    Which of the following statements accurately describe the Wav2Vec2 model and its components? (Select ALL that apply)
    1. [[IMAGE:0b0841408a2ed374_5_14]]
      Source diagram or notation
    2. [[IMAGE:0b0841408a2ed374_5_15]]
      Source diagram or notation
    3. [[IMAGE:0b0841408a2ed374_6_16]]
      Source diagram or notation
    4. [[IMAGE:0b0841408a2ed374_6_17]]
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 210 MSQ · 4.0 marks

    Regarding the SpeechT5 model for Text-to-Speech (TTS), which of the following statements are true? (Select ALL that apply)
    1. It is an encoder-decoder Transformer model.
    2. To generate speech for a specific voice, it requires a speaker embedding as an additional input.
    3. HiFi-GAN is an integrated and mandatory part of the SpeechT5 model architecture.
    4. It is pre-trained only on text data, similar to BERT.

    A published solution is not available for this question yet.

    Question 211 MSQ · 4.0 marks

    [[IMAGE:0b0841408a2ed374_6_18]]
    Source diagram or notation
    1. [[IMAGE:0b0841408a2ed374_6_19]]
      Source diagram or notation
    2. [[IMAGE:0b0841408a2ed374_6_20]]
      Source diagram or notation
    3. [[IMAGE:0b0841408a2ed374_6_21]]
      Source diagram or notation
    4. [[IMAGE:0b0841408a2ed374_6_22]]
      Source diagram or notation

    A published solution is not available for this question yet.

    Question 212 NAT · 3.0 marks

    You are evaluating an ASR model using 'jiwer'. The ground truth transcription is: "the quick brown fox jumps over the lazy dog" (9 words). The model predicts: "the quick fox jumped over a lazy dog". Calculate the Word Error Rate (WER). Provide the answer as a decimal rounded to two places.

      A published solution is not available for this question yet.

      Question 213 NAT · 3.0 marks

      A 3.5-second audio file is sampled at 16000 Hz. It is fed into a 'Wav2Vec2Model' ("facebook/wav2vec2-base-960h"), which uses a CNN feature extractor that downsamples the input by a factor of 320. What will be the sequence length (i.e., the number of time steps) of the 'last hidden state' output?

        A published solution is not available for this question yet.

        Question 214 NAT · 4.0 marks

        A speaker diarization script processes a dataset with 200 unique speakers. It extracts an average of 15 segments for each speaker using a model, which produces 512- dimensional embeddings. If each embedding vector is stored using 32-bit floating-point precision (4 bytes per value), what is the total memory required to store all the segment embeddings in Megabytes (MB)? (Assume 1 MB = 1024\(^{2}\) bytes). Round your answer to the nearest integer.

          A published solution is not available for this question yet.