MauryaHub PYQ Practice

da5013_2025T2_Q1_NA.pdf

Deep Learning Practice · Quiz 1 · May 2025

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 213 MCQ · 2.0 marks

A start-up is building a new language model for a low-resource language with many compound words and complex morphology. They are debating tokenization strategies. Which of the following approaches is most likely to offer the best balance between vocabulary size, handling OOV words, and capturing morphological variants effectively for this scenario?
  1. Character-level tokenization
  2. Word-level tokenization with a fixed vocabulary of 50,000 common words.
  3. Subword tokenization (e.g., BPE or SentencePiece) trained on the available corpus.
  4. Using only pre-defined special tokens and treating all other text as raw byte sequences.

A published solution is not available for this question yet.

Question 214 MCQ · 2.0 marks

When fully fine-tuning a large pre-trained Transformer model (e.g., >1 Billion parameters), which of the following contributes LEAST significantly to the GPU memory bottleneck compared to the others?
  1. Storing the model parameters themselves.
  2. Storing the gradients for each parameter.
  3. Storing the optimizer states (e.g., momentum and variance for Adam).
  4. Storing the input batch data (token IDs).

A published solution is not available for this question yet.

Question 215 MCQ · 2.0 marks

A research team wants their pre-trained language model to generate more helpful and harmless responses without extensive task-specific dataset collection. They have a collection of prompts and human-preferred responses. Which of the following techniques directly aligns with this goal and data?
  1. Pre-training the model on a larger, more diverse text corpus.
  2. Full fine-tuning on multiple downstream classification tasks.
  3. Instruction Tuning using prompt-completion pairs or reformatting existing datasets into an instructional format.
  4. Implementing Parameter-Efficient Fine-Tuning (PEFT) techniques like LoRa.

A published solution is not available for this question yet.

Question 216 MCQ · 3.0 marks

[[IMAGE:2b90cb5d8dd5ff90_4_0]]
Source diagram or notation
  1. [[IMAGE:2b90cb5d8dd5ff90_4_1]]
    Source diagram or notation
  2. [[IMAGE:2b90cb5d8dd5ff90_4_2]]
    Source diagram or notation
  3. [[IMAGE:2b90cb5d8dd5ff90_4_3]]
    Source diagram or notation
  4. [[IMAGE:2b90cb5d8dd5ff90_4_4]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 217 MCQ · 3.0 marks

[[IMAGE:2b90cb5d8dd5ff90_5_5]]
Source diagram or notation
  1. [[IMAGE:2b90cb5d8dd5ff90_5_6]]
    Source diagram or notation
  2. [[IMAGE:2b90cb5d8dd5ff90_5_7]]
    Source diagram or notation
  3. [[IMAGE:2b90cb5d8dd5ff90_5_8]]
    Source diagram or notation
  4. [[IMAGE:2b90cb5d8dd5ff90_5_9]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 218 MSQ · 4.0 marks

Which of the following statements accurately describe common characteristics or goals of subword tokenization algorithms like BPE, WordPiece, or SentencePiece? (Select ALL that apply)
  1. They aim to significantly reduce the vocabulary size compared to character- level tokenization.
  2. They can handle out-of-vocabulary (OOV) words by breaking them into known subword units.
  3. They primarily rely on merging the least frequent character or subword pairs to build the vocabulary.
  4. They can represent common words as single tokens and rare words as sequences of subword tokens.
  5. SentencePiece is designed to be language-agnostic, not requiring pre- segmentation based on spaces.

A published solution is not available for this question yet.

Question 219 MSQ · 4.0 marks

A team has a powerful pre-trained language model (e.g., a GPT-3 class model). They want to adapt it for a new summarization task but have very limited labeled summarization data and limited compute for full fine-tuning. Which of the following strategies could be viable and effective? (Select ALL that apply)
  1. Zero-shot prompting by providing the text and an instruction like "Summarize this:".
  2. Few-shot prompting (in-context learning) by providing a few examples of text and their summaries in the prompt before the target text.
  3. Full supervised fine-tuning of all model parameters on the small labeled dataset.
  4. Using a Parameter-Efficient Fine-Tuning (PEFT) method like LoRA on the small labeled dataset.
  5. Collecting a much larger unlabeled corpus related to the summarization domain and continuing pre-training.

A published solution is not available for this question yet.

Question 220 MSQ · 4.0 marks

[[IMAGE:2b90cb5d8dd5ff90_7_10]]
Source diagram or notation
  1. [[IMAGE:2b90cb5d8dd5ff90_7_11]]
    Source diagram or notation
  2. [[IMAGE:2b90cb5d8dd5ff90_7_12]]
    Source diagram or notation
  3. [[IMAGE:2b90cb5d8dd5ff90_7_13]]
    Source diagram or notation
  4. [[IMAGE:2b90cb5d8dd5ff90_7_14]]
    Source diagram or notation
  5. [[IMAGE:2b90cb5d8dd5ff90_7_15]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 221 NAT · 3.0 marks

[[IMAGE:2b90cb5d8dd5ff90_8_16]]
Source diagram or notation

    A published solution is not available for this question yet.

    Question 222 NAT · 3.0 marks

    [[IMAGE:2b90cb5d8dd5ff90_8_17]]
    Source diagram or notation

      A published solution is not available for this question yet.

      Question 223 NAT · 4.0 marks

      [[IMAGE:2b90cb5d8dd5ff90_9_18]]
      Source diagram or notation

        A published solution is not available for this question yet.

        Question 224 MCQ · 3.0 marks

        [[IMAGE:2b90cb5d8dd5ff90_10_19]] Based on the above data, answer the given subquestions.
        [[IMAGE:2b90cb5d8dd5ff90_10_20]]
        Source diagram or notationSource diagram or notation
        1. [[IMAGE:2b90cb5d8dd5ff90_10_21]]
          Source diagram or notation
        2. [[IMAGE:2b90cb5d8dd5ff90_10_22]]
          Source diagram or notation
        3. [[IMAGE:2b90cb5d8dd5ff90_10_23]]
          Source diagram or notation
        4. [[IMAGE:2b90cb5d8dd5ff90_10_24]]
          Source diagram or notation

        A published solution is not available for this question yet.

        Question 225 NAT · 4.0 marks

        [[IMAGE:2b90cb5d8dd5ff90_10_19]] Based on the above data, answer the given subquestions.
        Based on the provided configuration, calculate the total number of parameters in the model’s embedding layer (token embeddings) in millions. Enter your answer rounded to one decimal place.
        Source diagram or notation

          A published solution is not available for this question yet.

          Question 226 MCQ · 3.0 marks

          [[IMAGE:2b90cb5d8dd5ff90_10_19]] Based on the above data, answer the given subquestions.
          Based on the provided configuration, what is a primary characteristic of this language model’s architecture and training paradigm?
          Source diagram or notation
          1. It’s an encoder-decoder model designed for sequence-to-sequence tasks like translation.
          2. It’s an encoder-only model, likely using Masked Language Modeling for pre- training.
          3. It’s a decoder-only model, pre-trained using a Causal Language Modeling objective.
          4. It’s a small model primarily intended for edge devices due to its limited context length.

          A published solution is not available for this question yet.

          Question 227 NAT · 4.0 marks

          [[IMAGE:2b90cb5d8dd5ff90_10_19]] Based on the above data, answer the given subquestions.
          Considering the Adam optimizer stores 2 floating-point values per model parameter and parameters are 32-bit floats (4 bytes), if the total number of trainable parameters in the model is exactly 350 Million, how much GPU memory (in Gigabytes, GB) would be required just for the optimizer states? (Assume 1 GB = 10\(^{9}\) bytes). Enter your answer rounded to one decimal place.
          Source diagram or notation

            A published solution is not available for this question yet.

            Question 228 MCQ · 2.0 marks

            [[IMAGE:2b90cb5d8dd5ff90_10_19]] Based on the above data, answer the given subquestions.
            The configuration states the model uses Byte Pair Encoding (BPE). What is a key implication of this choice for handling text from diverse sources during inference?
            Source diagram or notation
            1. The model will be unable to process any words not seen during BPE vocabulary training, leading to frequent errors.
            2. All input words will be tokenized into individual characters, increasing sequence length significantly.
            3. Out-of-vocabulary words can be represented as sequences of known subword units, allowing the model to process them.
            4. BPE ensures that every language will have roughly the same number of tokens for a text of similar semantic content.       **Statistical Computing** **Section Id :** 64065391705 **Section Number :** 12 **Section type :** Online **Mandatory or Optional :** Mandatory **Number of Questions :** 10 **Number of Questions to be attempted :** 10 **Section Marks :** 30 **Display Number Panel :** Yes **Section Negative Marks :** 0 **Group All Questions :** No **Enable Mark as Answered Mark for Review and** No **Clear Response :** **Section Maximum Duration :** 0 **Section Minimum Duration :** 0 **Section Time In :** Minutes **Maximum Instruction Time :** 0

            A published solution is not available for this question yet.