MauryaHub PYQ Practice

da5013_2025T3_Q1_NA.pdf

Deep Learning Practice · Quiz 1 · Sep 2025

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 77 MCQ · 1.0 marks

Which one of the Nash equilibria is also a subgame perfect equilibrium?
  1. (AE, D)
  2. (AF, D)
  3. (BE, C)
  4. (BF, C)

A published solution is not available for this question yet.

Question 78 MCQ · 1.0 marks

No. of sub games this game has
  1. 1
  2. 2
  3. 3
  4. 4       **DLP** **Section Id :** 640653106432 **Section Number :** 5 **Section type :** Online **Mandatory or Optional :** Mandatory **Number of Questions :** 16 **Number of Questions to be attempted :** 16 **Section Marks :** 50 **Display Number Panel :** Yes **Section Negative Marks :** 0 **Group All Questions :** No **Enable Mark as Answered Mark for Review and** No **Clear Response :** **Section Maximum Duration :** 0 **Section Minimum Duration :** 0 **Section Time In :** Minutes **Maximum Instruction Time :** 0

A published solution is not available for this question yet.

Question 80 MCQ · 2.0 marks

Which of the following statements best describes the primary advantage of the SentencePiece tokenizer compared to a standard BPE implementation?
  1. It is significantly faster to train because it does not need to count pairs.
  2. It results in a smaller vocabulary size by always merging the shortest tokens first.
  3. It is inherently language-agnostic, treating text as a raw stream of Unicode characters, which is ideal for languages without clear word delimiters like Japanese or Thai.
  4. It is deterministic and always produces the same tokenization for a given string,unlike BPE which can be probabilistic.

A published solution is not available for this question yet.

Question 81 MCQ · 2.0 marks

When performing full fine-tuning of a large language model (e.g., 10B parameters) using Adam, which component consumes the most GPU memory?
  1. The model’s weights (parameters).
  2. The gradients calculated for each parameter.
  3. The optimizer states (e.g., momentum and variance).
  4. The vocabulary and embedding matrix.

A published solution is not available for this question yet.

Question 82 MCQ · 2.0 marks

A company wants to align its chatbot with values of being helpful, harmless, and honest. Human labelers provide ideal responses and rank AI-generated outputs. Which adaptation technique is designed for this?
  1. Instruction Tuning
  2. Zero-shot prompting
  3. Reinforcement Learning from Human Feedback (RLHF)
  4. Continued pre-training

A published solution is not available for this question yet.

Question 83 MCQ · 3.0 marks

[[IMAGE:8fbdb82708fcb40c_3_0]]
Source diagram or notation
  1. [[IMAGE:8fbdb82708fcb40c_3_1]]
    Source diagram or notation
  2. [[IMAGE:8fbdb82708fcb40c_3_2]]
    Source diagram or notation
  3. [[IMAGE:8fbdb82708fcb40c_3_3]]
    Source diagram or notation
  4. [[IMAGE:8fbdb82708fcb40c_3_4]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 84 MCQ · 3.0 marks

[[IMAGE:8fbdb82708fcb40c_4_5]]
Source diagram or notation
  1. [[IMAGE:8fbdb82708fcb40c_4_6]]
    Source diagram or notation
  2. [[IMAGE:8fbdb82708fcb40c_4_7]]
    Source diagram or notation
  3. [[IMAGE:8fbdb82708fcb40c_4_8]]
    Source diagram or notation
  4. [[IMAGE:8fbdb82708fcb40c_4_9]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 85 MCQ · 3.0 marks

What is a key implication of using a Parameter-Efficient Fine-Tuning (PEFT) method like LoRA when adapting a large language model for a new task?
  1. Inference speed is 10x faster.
  2. Only a small number of new parameters are trained while freezing original weights.
  3. It eliminates the need for labeled data.
  4. Model must be retrained from scratch.

A published solution is not available for this question yet.

Question 86 MCQ · 4.0 marks

The WordPiece tokenization algorithm, unlike BPE, does not merge the pair with the highest frequency. Instead, it merges the pair that maximizes a likelihood score. Which of the following statements accurately describes this score and its implication?
  1. [[IMAGE:8fbdb82708fcb40c_5_10]]
    Source diagram or notation
  2. [[IMAGE:8fbdb82708fcb40c_5_11]]
    Source diagram or notation
  3. [[IMAGE:8fbdb82708fcb40c_5_12]]
    Source diagram or notation
  4. [[IMAGE:8fbdb82708fcb40c_5_13]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 87 MSQ · 4.0 marks

Which of the following statements accurately describes the Causal Language Modeling (CLM) objective used to pre-train models like GPT? (Select ALL that apply)
  1. Predicting randomly masked tokens.
  2. Auto-regressive next-token prediction.
  3. Requires causal attention mask.
  4. Suited for encoder-only models.
  5. Maximizes joint probability of sequence.

A published solution is not available for this question yet.

Question 88 MSQ · 4.0 marks

A research lab has access to a powerful 175B parameter language model. They need to adapt it for a highly specialized legal text analysis task, but they only have about 500 labeled examples and limited access to high-end GPUs for fine-tuning. Which of the following are viable and computationally efficient adaptation strategies? (Select ALL that apply)
  1. Full fine-tuning
  2. Zero-shot prompting
  3. Few-shot in-context learning
  4. LoRA (PEFT)
  5. Pre-training from scratch

A published solution is not available for this question yet.

Question 89 MSQ · 4.0 marks

A team is fine-tuning a 7B parameter model on a single GPU with 24GB of memory. They are using the Adam optimizer (which stores 2 states per parameter) and 32-bit precision (4 bytes per parameter/ state/gradient). They find that they run out of memory even with a batch size of 1. Which of the following strategies could help them complete the fine-tuning process on this GPU? (Select ALL that apply)
  1. Increase learning rate
  2. Use LoRA
  3. Use quantization (8/4-bit)
  4. Switch to SGD
  5. Gradient accumulation

A published solution is not available for this question yet.

Question 90 MSQ · 4.0 marks

The evolution of NLP models shows a distinct shift from task-specific architectures to a ”pre-train, finetune” paradigm, and now towards large-scale, general-purpose models. Which of the following accurately represents this evolution and the capabilities at each stage? (Select ALL that apply)
  1. The earliest models (e.g., n-grams) were statistical, required task-specific design, and had limited generalization capacity.
  2. The ”pre-train, fine-tune” era (e.g., BERT, GPT) introduced transfer learning, where a model was first trained on a general language task and then fully adapted to a specific downstream task.
  3. Modern Large Language Models (LLMs like GPT-4) exhibit ”emerging abilities,” allowing them to perform new tasks with zero or few examples (in-context learning) without any weight updates.
  4. Word2vec was a complete language model capable of generating text, similar to GPT.
  5. The primary innovation of transformers over RNNs was the use of recurrent connections, which made them more efficient to train on parallel hardware.

A published solution is not available for this question yet.

Question 91 MSQ · 4.0 marks

The three main families of Transformer-based models are Encoder-only (e.g., BERT), Decoder-only (e.g.,GPT), and Encoder-Decoder (e.g., T5, BART). Match the architecture to its most suitable pre- training objective and typical use case. (Select ALL that apply)
  1. Encoder-only models are best for natural language understanding tasks (like sentiment classification) and are often pre-trained with a Masked Language Modeling (MLM) objective.
  2. Decoder-only models are ideal for text generation tasks and are pre-trained with a Causal Language Modeling (CLM) objective.
  3. Encoder-Decoder models are most suitable for sequence-to-sequence tasks like translation or summarization.
  4. All three architectures are pre-trained using the same Causal Language Modeling objective.

A published solution is not available for this question yet.

Question 92 NAT · 3.0 marks

[[IMAGE:8fbdb82708fcb40c_7_14]]
Source diagram or notation

    A published solution is not available for this question yet.

    Question 93 NAT · 4.0 marks

    A single transformer block in a GPT-style model has the following configuration: embedding dimension (d model) = 1024, num attention heads = 16. Each attention head has a dimension of d model /num attention heads. Calculate the total number of parameters (weights and biases) for the self attention mechanism (specifically the Q, K, V, and Output projection layers) within this single block. Report the answer in millions, rounded to one decimal place.(in M)

      A published solution is not available for this question yet.