MauryaHub PYQ Practice

da5004_2025T3_ET_FN.pdf

Large Language Models · End Term · Sep 2025 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 1.0 marks

What happens when you use Batch Normalization with a batch size of 1?
  1. It works perfectly fine
  2. The statistics become meaningless (mean=0, std=0)
  3. It automatically switches to Layer Normalization
  4. It uses the running statistics from training

A published solution is not available for this question yet.

Question 3 MCQ · 2.0 marks

A language model outputs the following logits for the next token: {cat: 3.2, dog: 2.9, bird: 1.1, fish: 0.5, snake: -0.4} Before sampling, the decoding pipeline performs: • Temperature scaling with T = 2.0 • Top-K filtering with K = 3 After applying both steps, which tokens remain eligible for sampling?
  1. Only ''cat''
  2. 'cat'' or ''dog''
  3. 'cat'', ''dog'' or ''bird''
  4. 'dog'', ''bird'' or ''fish''

A published solution is not available for this question yet.

Question 4 MCQ · 2.0 marks

If a model's next-token probabilities are [[IMAGE:b55b6f7161c9d349_2_2]] , [[IMAGE:b55b6f7161c9d349_2_3]] , [[IMAGE:b55b6f7161c9d349_2_4]] , [[IMAGE:b55b6f7161c9d349_2_5]] , and [[IMAGE:b55b6f7161c9d349_2_6]] , and Top-P is set to [[IMAGE:b55b6f7161c9d349_2_7]] . Which set of tokens will be included in the nucleus for sampling?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. Tokens A, B, C, D, E
  2. Tokens A, B
  3. Tokens A, B, C
  4. Tokens A, B, C, D
  5. Tokens A, B, D

A published solution is not available for this question yet.

Question 5 MCQ · 2.0 marks

In a WordPiece vocabulary-building step, the corpus contains the following tokenized sequences: • [[IMAGE:b55b6f7161c9d349_3_8]] • [[IMAGE:b55b6f7161c9d349_3_9]] The current vocabulary includes the tokens: [[IMAGE:b55b6f7161c9d349_3_10]] , [[IMAGE:b55b6f7161c9d349_3_11]] , [[IMAGE:b55b6f7161c9d349_3_12]] , [[IMAGE:b55b6f7161c9d349_3_13]] , [[IMAGE:b55b6f7161c9d349_3_14]] Two candidate merges are being evaluated: • Merge A: [[IMAGE:b55b6f7161c9d349_3_15]] [[IMAGE:b55b6f7161c9d349_3_16]] [[IMAGE:b55b6f7161c9d349_3_17]] • Merge B: [[IMAGE:b55b6f7161c9d349_3_18]] [[IMAGE:b55b6f7161c9d349_3_19]] [[IMAGE:b55b6f7161c9d349_3_20]] Based purely on WordPiece's likelihood-based merge selection, which merge is more likely to be chosen?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. Merge [[IMAGE:b55b6f7161c9d349_3_21]]
    Source diagram or notation
  2. Merge [[IMAGE:b55b6f7161c9d349_3_22]]
    Source diagram or notation
  3. Insufficient information to determine

A published solution is not available for this question yet.

Question 6 MCQ · 2.0 marks

A research team wants to classify scientific abstracts into multiple topics simultaneously (e.g., ''ML'', ''biology'', ''statistics''), where each abstract may belong to more than one topic. They fine- tune BERT for this task. Which modification is MOST appropriate?
  1. Replace the [CLS] vector with an average of the top-4 attention heads
  2. Feed the [CLS] embedding into a linear layer with a sigmoid activation per label
  3. Use token embeddings individually and classify each token
  4. Use the [SEP] token embedding for multi-label prediction

A published solution is not available for this question yet.

Question 7 MCQ · 1.0 marks

A company wants to automatically correct noisy OCR text extracted from scanned documents. The text contains spelling mistakes, missing words, and scrambled phrases. Which model should they fine-tune?
  1. BERT, because it masks tokens and predicts them independently
  2. BART, because it is trained with text corruption and autoregressive reconstruction
  3. GPT, because it is optimal for bidirectional correction
  4. BERT, because [CLS] captures global structure

A published solution is not available for this question yet.

Question 8 MCQ · 1.0 marks

Computing the full self-attention matrix has a time complexity of [[IMAGE:b55b6f7161c9d349_4_23]] where [[IMAGE:b55b6f7161c9d349_4_24]] is the sequence length and [[IMAGE:b55b6f7161c9d349_4_25]] is the embedding dimension. Which step in the self-attention mechanism is primarily responsible for this quadratic scaling?
Source diagram or notationSource diagram or notationSource diagram or notation
  1. Computing the softmax for each row
  2. Computing the matrix product [[IMAGE:b55b6f7161c9d349_4_26]]
    Source diagram or notation
  3. Multiplying the attention matrix [[IMAGE:b55b6f7161c9d349_4_27]] with the value matrix [[IMAGE:b55b6f7161c9d349_4_28]]
    Source diagram or notationSource diagram or notation
  4. Performing layer normalization

A published solution is not available for this question yet.

Question 9 MCQ · 1.0 marks

During autoregressive inference, key-value (KV) caching is used to avoid recomputing keys and values for previously generated tokens. If the KV cache is allowed to grow without any limit, which of the following failure modes can occur?
  1. GPU memory exhaustion
  2. Latency becoming quadratic in the output length
  3. Inability to perform parallel token generation
  4. Beam search collapsing to a single token

A published solution is not available for this question yet.

Question 10 MCQ · 2.0 marks

Consider a sequence of length [[IMAGE:b55b6f7161c9d349_5_29]] processed by a local [[IMAGE:b55b6f7161c9d349_5_30]] mechanism with block size [[IMAGE:b55b6f7161c9d349_5_31]] , where [[IMAGE:b55b6f7161c9d349_5_32]] is a multiple of [[IMAGE:b55b6f7161c9d349_5_33]] . The sequence is partitioned into [[IMAGE:b55b6f7161c9d349_5_34]] non-overlapping contiguous blocks, and self-attention is computed [[IMAGE:b55b6f7161c9d349_5_35]] (i.e., attention is restricted to the main diagonal blocks of the [[IMAGE:b55b6f7161c9d349_5_36]] attention matrix). Ignoring the cost of linear projections for [[IMAGE:b55b6f7161c9d349_5_37]] , what is the time complexity of the self- attention operation for one layer in terms of [[IMAGE:b55b6f7161c9d349_5_38]] , [[IMAGE:b55b6f7161c9d349_5_39]] , and the head dimension [[IMAGE:b55b6f7161c9d349_5_40]] ?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. [[IMAGE:b55b6f7161c9d349_5_41]]
    Source diagram or notation
  2. [[IMAGE:b55b6f7161c9d349_5_42]]
    Source diagram or notation
  3. [[IMAGE:b55b6f7161c9d349_5_43]]
    Source diagram or notation
  4. [[IMAGE:b55b6f7161c9d349_5_44]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 11 MCQ · 2.0 marks

A 2D vector [[IMAGE:b55b6f7161c9d349_5_45]] is rotated by [[IMAGE:b55b6f7161c9d349_5_46]] using RoPE. What is the resulting vector?
Source diagram or notationSource diagram or notation
  1. [[IMAGE:b55b6f7161c9d349_5_47]]
    Source diagram or notation
  2. [[IMAGE:b55b6f7161c9d349_5_48]]
    Source diagram or notation
  3. [[IMAGE:b55b6f7161c9d349_5_49]]
    Source diagram or notation
  4. [[IMAGE:b55b6f7161c9d349_5_50]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 12 MCQ · 2.0 marks

If a model uses Absolute Positional Embeddings (APE), increasing the sequence length from 100 to 1000 requires:
  1. No new parameters
  2. 900 additional positional embedding vectors
  3. A new rotation matrix
  4. A distance-based bias matrix

A published solution is not available for this question yet.

Question 13 MCQ · 3.0 marks

In a sequence of length [[IMAGE:b55b6f7161c9d349_6_51]] , how many different attention pairs produce a distance of [[IMAGE:b55b6f7161c9d349_6_52]] ?
Source diagram or notationSource diagram or notation
  1. 3
  2. 5
  3. 6
  4. 7

A published solution is not available for this question yet.

Question 14 MCQ · 3.0 marks

For a sequence of length [[IMAGE:b55b6f7161c9d349_6_53]] , how many learned positional vectors are required by the following methods? • Absolute Positional Encoding (APE) • Relative Positional Encoding (RPE) • ALiBi • NoPE (No Positional Encoding)
Source diagram or notation
  1. APE = 10, RPE = 19, ALiBi = 0, NoPE = 0
  2. APE = 10, RPE = 10, ALiBi = 19, NoPE = 0
  3. APE = 20, RPE = 19, ALiBi = 1, NoPE = 1
  4. APE = 10, RPE = 2T = 20, ALiBi = 10, NoPE = 0

A published solution is not available for this question yet.

Question 15 MCQ · 2.0 marks

Which of the following methods uses/use the concept: "The farther apart two tokens are, the less they should attend to each other''?
  1. RoPE
  2. APE
  3. ALiBi
  4. RPE

A published solution is not available for this question yet.

Question 16 MCQ · 2.0 marks

You trained a model with max sequence length 512. At inference, you need to process sequence length 4096 without retraining. Which positional encoding will perform best out-of-the-box?
  1. APE
  2. RPE
  3. RoPE
  4. ALiBi

A published solution is not available for this question yet.

Question 17 MCQ · 2.0 marks

A model using ALiBi is extended from context length 1K [[IMAGE:b55b6f7161c9d349_7_54]] 8K. How many new ALiBi parameters are added?
Source diagram or notation
  1. 7000
  2. 8000
  3. Depends on the number of distances
  4. 0

A published solution is not available for this question yet.

Question 18 NAT · 3.0 marks

A mini-batch has [[IMAGE:b55b6f7161c9d349_8_55]] sentences, each padded to a length [[IMAGE:b55b6f7161c9d349_8_56]] . Self-attention (single head) computes one score matrix per sentence. How many scalar scores are computed in total for a single head across the batch?
Source diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 19 NAT · 2.0 marks

    Suppose you are working on prefix language modeling. The sequence length is [[IMAGE:b55b6f7161c9d349_8_57]] and the first two tokens represent the task-specific prefix. How many non-infinity elements are there in the mask for computing attention scores?
    Source diagram or notation

      A published solution is not available for this question yet.

      Question 20 NAT · 2.0 marks

      A model uses a hybrid attention pattern in which each token attends to: • A local window of size [[IMAGE:b55b6f7161c9d349_9_58]] • A set of [[IMAGE:b55b6f7161c9d349_9_59]] random tokens, • A set of [[IMAGE:b55b6f7161c9d349_9_60]] global tokens. Compute the total number of attention computations per token (integer expression).
      Source diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 21 NAT · 2.0 marks

        During deployment of a transformer-based Large Language Model (LLM), you want to estimate the memory usage of the KV cache for efficient batching. Compute the KV-cache memory required per token (in Bytes) for the following model configuration: • Number of Transformer Blocks ( [[IMAGE:b55b6f7161c9d349_9_61]] ): 24 • Attention Heads per Layer ( [[IMAGE:b55b6f7161c9d349_9_62]] ): 12 • Head Dimension ( [[IMAGE:b55b6f7161c9d349_9_63]] ): 96 • Precision used for KV tensors: 2 Bytes (FP16)
        Source diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 22 NAT · 3.0 marks

          You are implementing a custom Block Sparse Attention kernel. Given a sequence length [[IMAGE:b55b6f7161c9d349_10_64]] and block size [[IMAGE:b55b6f7161c9d349_10_65]] , you decide to compute the main diagonal blocks plus one random off-diagonal block for each block-row. How many total [[IMAGE:b55b6f7161c9d349_10_66]] block computations are performed?
          Source diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 23 MSQ · 3.0 marks

            A GPT-style causal language model is trained using the next-token prediction objective: [[IMAGE:b55b6f7161c9d349_10_67]] During inference, the model must generate tokens autoregressively from left to right using only past context. Consider the following statements about GPT-style causal models.
            Source diagram or notation
            1. GPT cannot condition on future tokens during training because the causal mask zeros out all attention to positions [[IMAGE:b55b6f7161c9d349_10_68]]
              Source diagram or notation
            2. If two prefixes [[IMAGE:b55b6f7161c9d349_10_69]] and [[IMAGE:b55b6f7161c9d349_10_70]] have identical embeddings at all positions, GPT must assign identical next-token distributions for both contexts
              Source diagram or notationSource diagram or notation
            3. Removing positional encodings would force GPT to treat all permutations of the same set of tokens as equivalent contexts
            4. Causal masking ensures that the computational cost of training scales linearly with sequence length

            A published solution is not available for this question yet.

            Question 24 MSQ · 3.0 marks

            Under which of the following decoding settings can the model produce different outputs across multiple runs on the same prompt?
            1. Top-K sampling with [[IMAGE:b55b6f7161c9d349_11_71]] and temperature [[IMAGE:b55b6f7161c9d349_11_72]]
              Source diagram or notationSource diagram or notation
            2. Top-K sampling with [[IMAGE:b55b6f7161c9d349_11_73]] and temperature [[IMAGE:b55b6f7161c9d349_11_74]]
              Source diagram or notationSource diagram or notation
            3. Nucleus (Top-p) sampling with [[IMAGE:b55b6f7161c9d349_11_75]] and temperature [[IMAGE:b55b6f7161c9d349_11_76]]
              Source diagram or notationSource diagram or notation
            4. Greedy decoding after applying a temperature [[IMAGE:b55b6f7161c9d349_11_77]] to the logits
              Source diagram or notation
            5. Beam search with beam width [[IMAGE:b55b6f7161c9d349_11_78]]
              Source diagram or notation

            A published solution is not available for this question yet.

            Question 25 MSQ · 2.0 marks

            Consider a BART encoder with sequence length [[IMAGE:b55b6f7161c9d349_11_79]] and hidden size [[IMAGE:b55b6f7161c9d349_11_80]] (fixed). Now suppose the sequence length is doubled from [[IMAGE:b55b6f7161c9d349_11_81]] to [[IMAGE:b55b6f7161c9d349_11_82]] , keeping all other parameters unchanged. Which of the following statements are true?
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
            1. Self-attention FLOPs increase by a factor of [[IMAGE:b55b6f7161c9d349_11_83]] .
              Source diagram or notation
            2. Self-attention FLOPs increase by a factor of [[IMAGE:b55b6f7161c9d349_11_84]] .
              Source diagram or notation
            3. FFN FLOPs increase by a factor of [[IMAGE:b55b6f7161c9d349_11_85]] .
              Source diagram or notation
            4. FFN FLOPs increase by a factor of [[IMAGE:b55b6f7161c9d349_11_86]] .
              Source diagram or notation

            A published solution is not available for this question yet.