MauryaHub PYQ Practice

da5013_2026T1_Q1_NA.pdf

Deep Learning Practice · Quiz 1 · Jan 2026

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 2.0 marks

Consider subword tokenization methods such as Byte Pair Encoding (BPE) and WordPiece as employed in transformer-based language modeling pipelines. Which of the following statements correctly describes their fundamental properties?
  1. They always split words at true morpheme boundaries (the smallest units of meaning, e.g., “un-” + “break” + “able”).
  2. They learn merge rules from corpus statistics (frequency/likelihood), not token semantics
  3. They eliminate (Out Of Vocabulary) OOV by ensuring every word is a single token
  4. They generalize perfectly to out-of-domain text without increasing sequence length.

A published solution is not available for this question yet.

Question 3 MCQ · 2.0 marks

Consider two subword tokenizers trained on the same text corpus. Tokenizer A uses a vocabulary of 8,000 tokens, whereas Tokenizer B uses a vocabulary of 64,000 tokens. Which of the following statements most accurately characterizes the implications of these vocabulary sizes for training efficiency in transformer-based language models?
  1. Larger vocabulary always reduces sequence length and total compute.
  2. Smaller vocabulary always improves semantic alignment.
  3. Larger vocabulary reduces sequence length but increases embedding parameters.
  4. Vocabulary size has no effect once model is pretrained.

A published solution is not available for this question yet.

Question 4 MCQ · 2.0 marks

A subword tokenizer decomposes rare chemical entity names into a large number of fragments. During downstream fine-tuning, the model exhibits degraded performance on a chemical named entity recognition (NER) task. Which of the following represents the most principled corrective action?
  1. Increase the dropout rate during fine-tuning.
  2. Retrain the tokenizer using an in-domain (chemical) corpus.
  3. Reduce the batch size during training.
  4. Apply label smoothing to the loss function.

A published solution is not available for this question yet.

Question 5 MCQ · 2.0 marks

[[IMAGE:5098e0c90adb7cd8_4_2]]
Source diagram or notation
  1. The output sequence will always contain exactly 12 non-padding tokens.
  2. Special tokens may consume a portion of the 12-token length budget.
  3. Word boundaries are preserved, since truncation occurs at whitespace boundaries.
  4. Padding will never be applied when truncation is enabled.

A published solution is not available for this question yet.

Question 6 MCQ · 2.0 marks

[[IMAGE:5098e0c90adb7cd8_5_3]]
Source diagram or notation
  1. This is supervised fine-tuning with noisy labels.
  2. This is instruction tuning because labels are present.
  3. This performs continual pretraining with MLM objective.
  4. This trains a classifier head implicitly.

A published solution is not available for this question yet.

Question 7 MCQ · 2.0 marks

In the standard Transformer architecture (e.g., GPT or BERT), which component exhibits a computational complexity that scales quadratically ( [[IMAGE:5098e0c90adb7cd8_5_4]] ) with respect to the input sequence length [[IMAGE:5098e0c90adb7cd8_5_5]] ?
Source diagram or notationSource diagram or notation
  1. The initial token and positional embedding lookup.
  2. The element-wise addition in the Residual (Skip) connections.
  3. The computation of the attention score matrix ( [[IMAGE:5098e0c90adb7cd8_6_6]] ).
    Source diagram or notation
  4. The linear projections in the Position-wise Feed-Forward Network (FFN).

A published solution is not available for this question yet.

Question 8 MCQ · 2.0 marks

A model with 4B parameters (fp32) is being trained using the AdamW optimizer on a single 32GB GPU. Why is full fine-tuning impossible in this configuration?
  1. The 4B parameters alone require 32GB of VRAM, leaving no room for the OS or CUDA kernels.
  2. The optimizer states (32GB) plus the model weights (16GB) and gradients (16GB) exceed the 32GB VRAM limit.
  3. The tokenizer is unable to map a context length of 2048 to a 32-bit integer space.
  4. The fp32 precision requires 8 bytes per parameter, meaning the 4B model needs 64GB just to load.

A published solution is not available for this question yet.

Question 9 MCQ · 2.0 marks

During the full fine-tuning of a Large Language Model, you aim to regularize the training process by penalizing the growth of weight magnitudes. This ensures the model does not overfit the fine- tuning data by making excessively large updates to the pretrained parameters. Which parameter is specifically designed to apply this penalty?
  1. Gradient Accumulation.
  2. Batch Size.
  3. Checkpointing.
  4. Weight Decay.

A published solution is not available for this question yet.

Question 10 MCQ · 2.0 marks

Why does full fine-tuning of large language models often require more memory than inference using the same model?
  1. Inference uses larger batch sizes.
  2. Inference stores optimizer states.
  3. Training requires storing gradients and optimizer states.
  4. Training uses longer input sequences.

A published solution is not available for this question yet.

Question 11 MCQ · 2.0 marks

Which technique is most effective at reducing GPU memory usage during training without modifying the model architecture?
  1. Gradient clipping.
  2. Mixed-precision training.
  3. Increasing learning rate warmup.
  4. Saving checkpoints less frequently.

A published solution is not available for this question yet.

Question 12 MSQ · 3.0 marks

Consider the output of a Hugging Face tokenizer when applied to a batch of sequences with padding= "max_length" and return_tensors= "pt". Which of the following statements are correct? (Select ALL that apply)
  1. The input_ids tensor contains the numerical indices mapped from the vocabulary.
  2. The attention_mask contains 0s for padding tokens and 1s for real tokens to prevent the model from attending to padding.
  3. The token_type_ids (if present) are used primarily to distinguish between uppercase and lowercase letters.
  4. In models like BERT, the input_ids will typically begin with a special token index (e.g., [CLS]).

A published solution is not available for this question yet.

Question 13 MSQ · 3.0 marks

[[IMAGE:5098e0c90adb7cd8_8_7]]
Source diagram or notation
  1. Effective batch size is larger than 1.
  2. Reduced memory usage compared to FP32 training.
  3. Gradient clipping limits large parameter updates.
  4. Gradient accumulation reduces the number of optimizer states.

A published solution is not available for this question yet.

Question 14 MSQ · 3.0 marks

Which of the following statements correctly distinguish different Transformer architectures and their typical training objectives? (Select ALL that apply)
  1. Encoder-only models are commonly trained with Masked Language Modeling objectives.
  2. Decoder-only models rely on causal masking and autoregressive inference.
  3. Encoder-decoder models are well suited for sequence-to-sequence tasks.
  4. All Transformer architectures use identical attention masks.

A published solution is not available for this question yet.

Question 15 MSQ · 3.0 marks

[[IMAGE:5098e0c90adb7cd8_9_8]]
Source diagram or notation
  1. Task Conditioning: The model learns to follow natural language formatting (Instruction/Input/Answer) rather than just mapping a sequence to a single integer ID.
  2. Open Vocabulary: The model's prediction head remains the size of its full vocabulary (e.g., 50k+ tokens), allowing it to generate any string as an answer instead of being restricted to fixed logits for [[IMAGE:5098e0c90adb7cd8_9_9]] classes.
    Source diagram or notation
  3. Guaranteed Zero-Shot: This training format ensures the model will generalize perfectly to any unseen task prompt without further data.
  4. Objective Shift: The training objective moves from minimizing cross-entropy loss over a discrete class index to minimizing next-token prediction loss over the sequence tokens.

A published solution is not available for this question yet.

Question 16 MSQ · 3.0 marks

A model is trained using the Causal Language Modeling (CLM) objective with a standard cross- entropy loss. Which of the following statements correctly describe the training dynamics? (Select ALL that apply)
  1. The loss at position [[IMAGE:5098e0c90adb7cd8_9_10]] depends only on tokens [[IMAGE:5098e0c90adb7cd8_9_11]] .
    Source diagram or notationSource diagram or notation
  2. The attention mask is strictly upper triangular.
  3. Future tokens contribute gradients to earlier positions.
  4. The joint probability of the sequence is factorized autoregressively.

A published solution is not available for this question yet.

Question 17 MSQ · 3.0 marks

When configuring TrainingArguments for a Transformer model, we include a learning rate warmup phase (e.g., warmup_steps=500). Which of the following statements correctly describe the purpose and behavior of this strategy? (Select ALL that apply)
  1. It involves linearly increasing the learning rate from 0 (or a small value) to the target maximum during the initial phase.
  2. It helps prevent "divergence'' or instability caused by large gradients when the model weights are far from their optimal state.
  3. It significantly reduces the VRAM (memory) required to store optimizer states.
  4. After the warmup phase, the learning rate typically follows a decay schedule (like linear or cosine) to ensure convergence.

A published solution is not available for this question yet.

Question 18 NAT · 4.0 marks

A transformer-based language model has 1.2 billion parameters and is fine-tuned using AdamW. Assume parameters, gradients, and optimizer states (first and second moments) are all stored in 32-bit precision (4 bytes each). Ignoring activations and buffers, calculate the total GPU memory required (in GB) to store parameters, gradients, and optimizer states. (Assume 1 GB = [[IMAGE:5098e0c90adb7cd8_10_12]] bytes.)
Source diagram or notation

    A published solution is not available for this question yet.

    Question 19 NAT · 4.0 marks

    A GPT-style transformer block has an embedding dimension of 768 and uses 12 attention heads. Consider a specific processing task with a very short sequence of only 3 tokens ( [[IMAGE:5098e0c90adb7cd8_11_13]] ).Compute the total number of weight parameters (ignore biases) in the self-attention module, specifically for the Query, Key, Value, and Output projection matrices. Report your answer in millions, rounded to one decimal place.
    Source diagram or notation

      A published solution is not available for this question yet.

      Question 20 NAT · 4.0 marks

      You are training a decoder-only language model with a vocabulary size of 32,000, maximum context length of 2048, and embedding dimension of 1024. The model uses learned positional embeddings. Calculate the total number of parameters in the embedding layer (token + positional embeddings). Report your answer in millions, rounded to one decimal place.

        A published solution is not available for this question yet.