da5013_2025T3_Q1_NA.pdf
Deep Learning Practice · Quiz 1 · Sep 2025
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 77 MCQ · 1.0 marks
Which one of the Nash equilibria is also a subgame perfect equilibrium?
(AE, D)
(AF, D)
(BE, C)
(BF, C)
A published solution is not available for this question yet.
Question 78 MCQ · 1.0 marks
No. of sub games this game has
1
2
3
4
**DLP**
**Section Id :** 640653106432
**Section Number :** 5
**Section type :** Online
**Mandatory or Optional :** Mandatory
**Number of Questions :** 16
**Number of Questions to be attempted :** 16
**Section Marks :** 50
**Display Number Panel :** Yes
**Section Negative Marks :** 0
**Group All Questions :** No
**Enable Mark as Answered Mark for Review and**
No
**Clear Response :**
**Section Maximum Duration :** 0
**Section Minimum Duration :** 0
**Section Time In :** Minutes
**Maximum Instruction Time :** 0
A published solution is not available for this question yet.
Question 80 MCQ · 2.0 marks
Which of the following statements best describes the primary advantage of the SentencePiece
tokenizer compared to a standard BPE implementation?
It is significantly faster to train because it does not need to count pairs.
It results in a smaller vocabulary size by always merging the shortest tokens
first.
It is inherently language-agnostic, treating text as a raw stream of Unicode
characters, which is ideal for languages without clear word delimiters like Japanese or Thai.
It is deterministic and always produces the same tokenization for a given
string,unlike BPE which can be probabilistic.
A published solution is not available for this question yet.
Question 81 MCQ · 2.0 marks
When performing full fine-tuning of a large language model (e.g., 10B parameters) using Adam,
which component consumes the most GPU memory?
The model’s weights (parameters).
The gradients calculated for each parameter.
The optimizer states (e.g., momentum and variance).
The vocabulary and embedding matrix.
A published solution is not available for this question yet.
Question 82 MCQ · 2.0 marks
A company wants to align its chatbot with values of being helpful, harmless, and honest. Human
labelers provide ideal responses and rank AI-generated outputs. Which adaptation technique is
designed for this?
Instruction Tuning
Zero-shot prompting
Reinforcement Learning from Human Feedback (RLHF)
Continued pre-training
A published solution is not available for this question yet.
Question 83 MCQ · 3.0 marks
[[IMAGE:8fbdb82708fcb40c_3_0]]

[[IMAGE:8fbdb82708fcb40c_3_1]]

[[IMAGE:8fbdb82708fcb40c_3_2]]

[[IMAGE:8fbdb82708fcb40c_3_3]]

[[IMAGE:8fbdb82708fcb40c_3_4]]

A published solution is not available for this question yet.
Question 84 MCQ · 3.0 marks
[[IMAGE:8fbdb82708fcb40c_4_5]]

[[IMAGE:8fbdb82708fcb40c_4_6]]

[[IMAGE:8fbdb82708fcb40c_4_7]]

[[IMAGE:8fbdb82708fcb40c_4_8]]

[[IMAGE:8fbdb82708fcb40c_4_9]]

A published solution is not available for this question yet.
Question 85 MCQ · 3.0 marks
What is a key implication of using a Parameter-Efficient Fine-Tuning (PEFT) method like LoRA when
adapting a large language model for a new task?
Inference speed is 10x faster.
Only a small number of new parameters are trained while freezing original
weights.
It eliminates the need for labeled data.
Model must be retrained from scratch.
A published solution is not available for this question yet.
Question 86 MCQ · 4.0 marks
The WordPiece tokenization algorithm, unlike BPE, does not merge the pair with the highest
frequency. Instead, it merges the pair that maximizes a likelihood score. Which of the following
statements accurately describes this score and its implication?
[[IMAGE:8fbdb82708fcb40c_5_10]]

[[IMAGE:8fbdb82708fcb40c_5_11]]

[[IMAGE:8fbdb82708fcb40c_5_12]]

[[IMAGE:8fbdb82708fcb40c_5_13]]

A published solution is not available for this question yet.
Question 87 MSQ · 4.0 marks
Which of the following statements accurately describes the Causal Language Modeling (CLM)
objective used to pre-train models like GPT? (Select ALL that apply)
Predicting randomly masked tokens.
Auto-regressive next-token prediction.
Requires causal attention mask.
Suited for encoder-only models.
Maximizes joint probability of sequence.
A published solution is not available for this question yet.
Question 88 MSQ · 4.0 marks
A research lab has access to a powerful 175B parameter language model. They need to adapt it for
a highly specialized legal text analysis task, but they only have about 500 labeled examples and
limited access to high-end GPUs for fine-tuning. Which of the following are viable and
computationally efficient adaptation strategies? (Select ALL that apply)
Full fine-tuning
Zero-shot prompting
Few-shot in-context learning
LoRA (PEFT)
Pre-training from scratch
A published solution is not available for this question yet.
Question 89 MSQ · 4.0 marks
A team is fine-tuning a 7B parameter model on a single GPU with 24GB of memory. They are using
the Adam optimizer (which stores 2 states per parameter) and 32-bit precision (4 bytes per
parameter/ state/gradient). They find that they run out of memory even with a batch size of 1.
Which of the following strategies could help them complete the fine-tuning process on this GPU?
(Select ALL that apply)
Increase learning rate
Use LoRA
Use quantization (8/4-bit)
Switch to SGD
Gradient accumulation
A published solution is not available for this question yet.
Question 90 MSQ · 4.0 marks
The evolution of NLP models shows a distinct shift from task-specific architectures to a ”pre-train,
finetune” paradigm, and now towards large-scale, general-purpose models. Which of the following
accurately represents this evolution and the capabilities at each stage? (Select ALL that apply)
The earliest models (e.g., n-grams) were statistical, required task-specific
design, and had limited generalization capacity.
The ”pre-train, fine-tune” era (e.g., BERT, GPT) introduced transfer learning,
where a model was first trained on a general language task and then fully adapted to a specific
downstream task.
Modern Large Language Models (LLMs like GPT-4) exhibit ”emerging abilities,”
allowing them to perform new tasks with zero or few examples (in-context learning) without any
weight updates.
Word2vec was a complete language model capable of generating text, similar
to GPT.
The primary innovation of transformers over RNNs was the use of recurrent
connections, which made them more efficient to train on parallel hardware.
A published solution is not available for this question yet.
Question 91 MSQ · 4.0 marks
The three main families of Transformer-based models are Encoder-only (e.g., BERT), Decoder-only
(e.g.,GPT), and Encoder-Decoder (e.g., T5, BART). Match the architecture to its most suitable pre-
training objective and typical use case. (Select ALL that apply)
Encoder-only models are best for natural language understanding tasks (like
sentiment classification) and are often pre-trained with a Masked Language Modeling (MLM)
objective.
Decoder-only models are ideal for text generation tasks and are pre-trained
with a Causal Language Modeling (CLM) objective.
Encoder-Decoder models are most suitable for sequence-to-sequence tasks
like translation or summarization.
All three architectures are pre-trained using the same Causal Language
Modeling objective.
A published solution is not available for this question yet.
Question 92 NAT · 3.0 marks
[[IMAGE:8fbdb82708fcb40c_7_14]]

A published solution is not available for this question yet.
Question 93 NAT · 4.0 marks
A single transformer block in a GPT-style model has the following configuration: embedding
dimension (d model) = 1024, num attention heads = 16. Each attention head has a dimension of d
model /num attention heads. Calculate the total number of parameters (weights and biases) for
the self attention mechanism (specifically the Q, K, V, and Output projection layers) within this
single block. Report the answer in millions, rounded to one decimal place.(in M)
A published solution is not available for this question yet.