MauryaHub PYQ Practice

da5013_2026T2_Q1_NA.pdf

Deep Learning Practice · Quiz 1 · May 2026

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 2.0 marks

A startup is building an LLM for a low resource agglutinative language where a single word can contain information equivalent to an entire English sentence. After training a Word Level tokenizer with a vocabulary of 30,000, they observe that nearly 18% of tokens become [UNK]. Which action is most likely to improve the situation while keeping vocabulary growth manageable?
  1. Use SentencePiece or BPE
  2. Increase batch size
  3. Increase the minimum token frequency threshold during tokenizer training
  4. Disable normalization so every word form is stored separately

A published solution is not available for this question yet.

Question 3 MCQ · 2.0 marks

Consider the following code: [[IMAGE:be50b0cf89de8deb_3_2]] Assume that the variable [[IMAGE:be50b0cf89de8deb_3_3]] contains a piece of text. When tokenized without applying truncation or padding, it produces 220 tokens. Which of the following correctly describes the [[IMAGE:be50b0cf89de8deb_3_4]] returned in [[IMAGE:be50b0cf89de8deb_3_5]] ?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. The output contains all 220 tokens because tokenization is completed before truncation.
  2. The output contains the first 128 tokens followed by additional padding tokens.
  3. The output contains exactly 128 tokens from the document.
  4. The output length depends on the tokenizer model being used.

A published solution is not available for this question yet.

Question 4 MCQ · 2.0 marks

A team fine tunes a 13B parameter model using LoRA. After training they discover that only 0.15% of parameters have been updated. Which statement best explains this observation?
  1. LoRA updates low rank matrices while freezing most pretrained weights
  2. LoRA trains only the embedding layer
  3. LoRA automatically quantizes the model
  4. LoRA removes attention layers during training

A published solution is not available for this question yet.

Question 5 MCQ · 2.0 marks

A company wants to build a medical chatbot. They possess: • 8 million unlabeled medical documents • 300 doctor reviewed question answer pairs Which adaptation strategy would most likely provide the largest performance gain?
  1. Train a new tokenizer only
  2. Continue pretraining on the medical corpus and then instruction tune
  3. Increase context length
  4. Increase batch size

A published solution is not available for this question yet.

Question 6 MCQ · 2.0 marks

A 7 billion parameter model is loaded in FP32 precision. Ignoring activations and optimizer states, approximately how much GPU memory is required for storing model weights?
  1. 7 GB
  2. 14 GB
  3. 28 GB
  4. 56 GB

A published solution is not available for this question yet.

Question 7 MCQ · 2.0 marks

A tokenizer vocabulary is increased from 16K to 128K while keeping the training corpus unchanged. What tradeoff is most likely?
  1. Shorter sequences and a larger embedding layer
  2. Longer sequences and a smaller embedding layer
  3. Shorter sequences and a smaller embedding layer
  4. No significant effect on the model

A published solution is not available for this question yet.

Question 8 MCQ · 2.0 marks

Consider the following code: [[IMAGE:be50b0cf89de8deb_5_6]] Which method fills the blank to return **only** the token IDs as a plain Python list?
Source diagram or notation
  1. [[IMAGE:be50b0cf89de8deb_5_7]]
    Source diagram or notation
  2. [[IMAGE:be50b0cf89de8deb_5_8]]
    Source diagram or notation
  3. [[IMAGE:be50b0cf89de8deb_5_9]]
    Source diagram or notation
  4. [[IMAGE:be50b0cf89de8deb_5_10]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 9 MCQ · 3.0 marks

The following code is intended to implement causal masking for self attention, where each token can attend only to itself and previous tokens. [[IMAGE:be50b0cf89de8deb_6_11]] What is the effect of running this code as written? (**Note:** Assume necessary imports are already done and [[IMAGE:be50b0cf89de8deb_6_12]] is a valid tensor of attention scores with shape [[IMAGE:be50b0cf89de8deb_6_13]] .)
Source diagram or notationSource diagram or notationSource diagram or notation
  1. No effect. This correctly implements causal masking.
  2. It raises a [[IMAGE:be50b0cf89de8deb_6_14]] because [[IMAGE:be50b0cf89de8deb_6_15]] cannot be used with [[IMAGE:be50b0cf89de8deb_6_16]] .
    Source diagram or notationSource diagram or notationSource diagram or notation
  3. The attention weights in every row will sum to more than 1.
  4. The current token is also masked, so each token can attend only to earlier tokens. The first row becomes entirely [[IMAGE:be50b0cf89de8deb_6_17]] , which can lead to invalid ( [[IMAGE:be50b0cf89de8deb_6_18]] ) attention weights.
    Source diagram or notationSource diagram or notation

A published solution is not available for this question yet.

Question 10 MSQ · 3.0 marks

Consider the following training configuration: [[IMAGE:be50b0cf89de8deb_6_19]] Which statements are correct?
Source diagram or notation
  1. Effective batch size is larger than 2
  2. Memory usage is lower than FP32 training
  3. Optimizer updates occur after every batch
  4. Gradient accumulation can simulate larger batches
  5. FP16 automatically reduces optimizer state size by half

A published solution is not available for this question yet.

Question 11 MSQ · 3.0 marks

A student trains GPT 2 using Hugging Face [[IMAGE:be50b0cf89de8deb_7_20]] with: [[IMAGE:be50b0cf89de8deb_7_21]] The training dataset contains sequences of different lengths, but the student forgets to execute: [[IMAGE:be50b0cf89de8deb_7_22]] Which of the following are true?
Source diagram or notationSource diagram or notationSource diagram or notation
  1. The data collator will fail when it attempts to pad the batch.
  2. Decoder only models never need padding during training.
  3. Using [[IMAGE:be50b0cf89de8deb_7_23]] instead may cause padding to use a real vocabulary token.
    Source diagram or notation
  4. Reusing the EOS token as the padding token is a common practice for GPT 2.
  5. GPT 2 already defines a [[IMAGE:be50b0cf89de8deb_7_24]] token by default.
    Source diagram or notation

A published solution is not available for this question yet.

Question 12 MSQ · 3.0 marks

A team experiences GPU out of memory errors while fine tuning a language model. Which actions may help?
  1. Mixed precision training
  2. Quantization
  3. Gradient checkpointing
  4. LoRA
  5. Increasing vocabulary size

A published solution is not available for this question yet.

Question 13 MSQ · 3.0 marks

A model is trained for sentiment classification using the following prompt format: [[IMAGE:be50b0cf89de8deb_8_25]] Compared with traditional classification heads, which statements are true?
Source diagram or notation
  1. The model learns to follow instructions
  2. The output vocabulary remains open ended
  3. Predictions are restricted to predefined class IDs
  4. Training becomes a next token prediction task
  5. The objective is identical to logistic regression

A published solution is not available for this question yet.

Question 14 MSQ · 2.0 marks

Which outputs can typically be produced by a Hugging Face tokenizer?
  1. input_ids
  2. attention_mask
  3. token_type_ids
  4. hidden_states
  5. offset_mapping

A published solution is not available for this question yet.

Question 15 NAT · 4.0 marks

A GPT style model has the following configuration: • Vocabulary size = 65,000 • Context length = 4096 • Embedding dimension = 1536 The model uses learned positional embeddings. Calculate the total number of embedding parameters (token embeddings + positional embeddings) in millions. Round your answer to one decimal place.

    A published solution is not available for this question yet.

    Question 16 NAT · 4.0 marks

    A transformer block has: • d_model = 1536 Assume standard self attention consisting of: • Query projection • Key projection • Value projection • Output projection Ignore biases. Calculate the total number of parameters in millions. Round your answer to one decimal place.

      A published solution is not available for this question yet.

      Question 17 NAT · 4.0 marks

      A model contains 5 billion parameters. Training uses AdamW in FP32 precision. Assume: • Parameters = 4 bytes • Gradients = 4 bytes • Momentum = 4 bytes • Variance = 4 bytes Ignoring activations and buffers, calculate the total memory requirement in GB. Use: 1 GB = 10\(^{9}\) bytes Round your answer to one decimal place.

        A published solution is not available for this question yet.

        Question 18 NAT · 4.0 marks

        A batch of sentences is tokenized using a pretrained BERT tokenizer: [[IMAGE:be50b0cf89de8deb_11_26]] What is the value printed by the code? (Assume default tokenizer behavior is used.)
        Source diagram or notation

          A published solution is not available for this question yet.

          Question 19 NAT · 3.0 marks

          A training job uses: • Batch size = 8 • Context length = 4096 • Gradient accumulation = 4 • Training steps = 10,000 optimizer updates Calculate the total number of tokens processed after 10,000 optimizer updates. Express your answer in millions. Round to one decimal place.

            A published solution is not available for this question yet.