MauryaHub PYQ Practice

da5002_2026T2_ET_FN.pdf

Mathematical Foundations of Generative AI · End Term · May 2026 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 NAT · 2.0 marks

An agent receives rewards [[IMAGE:173306d6b8cbd418_2_2]] over an episode ( [[IMAGE:173306d6b8cbd418_2_3]] ) with the discount factor equal to 0.5. Compute the full discounted return [[IMAGE:173306d6b8cbd418_2_4]] for the trajectory [[IMAGE:173306d6b8cbd418_2_5]] . Enter your answer correct to three decimal places.
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 3 MSQ · 3.0 marks

    [[IMAGE:173306d6b8cbd418_3_6]]
    Source diagram or notation
    1. Both TRPO and PPO aim to prevent the updated policy from deviating too far from [[IMAGE:173306d6b8cbd418_3_7]] .
      Source diagram or notation
    2. PPO always produces a smaller policy update than TRPO for the same [[IMAGE:173306d6b8cbd418_3_8]] and [[IMAGE:173306d6b8cbd418_3_9]] .
      Source diagram or notationSource diagram or notation
    3. When [[IMAGE:173306d6b8cbd418_3_10]] is within [[IMAGE:173306d6b8cbd418_3_11]] , the PPO clip is inactive and [[IMAGE:173306d6b8cbd418_3_12]] reduces to the unclipped surrogate [[IMAGE:173306d6b8cbd418_3_13]] .
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
    4. Computing the PPO objective requires solving a constrained optimisation subproblem at each step, making it more expensive than TRPO.

    A published solution is not available for this question yet.

    Question 4 MSQ · 3.0 marks

    In Proximal Policy Optimization (PPO), the clipped surrogate objective for a single state-action pair is: [[IMAGE:173306d6b8cbd418_3_14]] Which of the following statements are correct?
    Source diagram or notation
    1. When [[IMAGE:173306d6b8cbd418_3_15]] and [[IMAGE:173306d6b8cbd418_3_16]] , the clipped objective prevents the gradient from further increasing [[IMAGE:173306d6b8cbd418_3_17]] , effectively imposing a trust region.
      Source diagram or notationSource diagram or notationSource diagram or notation
    2. When [[IMAGE:173306d6b8cbd418_4_18]] and [[IMAGE:173306d6b8cbd418_4_19]] , the clipped objective prevents the gradient from further decreasing [[IMAGE:173306d6b8cbd418_4_20]] .
      Source diagram or notationSource diagram or notationSource diagram or notation
    3. The clipping mechanism is equivalent to adding an explicit KL-divergence penalty [[IMAGE:173306d6b8cbd418_4_21]] to the objective.
      Source diagram or notation
    4. If [[IMAGE:173306d6b8cbd418_4_22]] , then [[IMAGE:173306d6b8cbd418_4_23]] regardless of the sign of [[IMAGE:173306d6b8cbd418_4_24]] .
      Source diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 5 MSQ · 3.0 marks

    In multi-head self-attention with [[IMAGE:173306d6b8cbd418_4_25]] heads, the queries, keys, and values for each head have dimension [[IMAGE:173306d6b8cbd418_4_26]] where [[IMAGE:173306d6b8cbd418_4_27]] is the model dimension. Which of the following correctly explain why this choice makes the implementation efficient?
    Source diagram or notationSource diagram or notationSource diagram or notation
    1. Choosing [[IMAGE:173306d6b8cbd418_4_28]] ensures that every head learns the same attention pattern while reducing the number of parameters.
      Source diagram or notation
    2. The total parameters across [[IMAGE:173306d6b8cbd418_4_29]] query projections [[IMAGE:173306d6b8cbd418_4_30]] equal [[IMAGE:173306d6b8cbd418_4_31]] , the same as a single-head projection [[IMAGE:173306d6b8cbd418_4_32]] . The same holds for key and value projections.
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
    3. Since [[IMAGE:173306d6b8cbd418_4_33]] , all [[IMAGE:173306d6b8cbd418_4_34]] heads can be computed as a single batched matrix multiplication, enabling parallel execution across heads on a GPU.
      Source diagram or notationSource diagram or notation
    4. Using [[IMAGE:173306d6b8cbd418_4_35]] , the total computational cost across all [[IMAGE:173306d6b8cbd418_4_36]] heads is reduced by a factor of [[IMAGE:173306d6b8cbd418_4_37]] compared to single-head attention with [[IMAGE:173306d6b8cbd418_4_38]] .
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 6 NAT · 3.0 marks

    [[IMAGE:173306d6b8cbd418_5_39]]
    Source diagram or notation

      A published solution is not available for this question yet.

      Question 7 MCQ · 2.0 marks

      What is the role of layer normalization in a transformer block?
      1. It eliminates the need for residual connections.
      2. It normalizes the hidden representation across its features for each token, helping maintain a stable scale of activations during training.
      3. It normalizes the hidden representations across different examples in a mini- batch.
      4. It introduces the nonlinearity needed by the transformer through a learnable activation function.

      A published solution is not available for this question yet.

      Question 8 MCQ · 2.0 marks

      Which among the following is true for a VAE and a diffusion model?
      1. Both models use an encoder and decoder to map between the input and a low-dimensional representation.
      2. Both the encoder and decoder are learned in VAE and diffusion.
      3. The hidden variables in a diffusion model are deterministic, whereas the latent variable in a VAE is stochastic.
      4. Both VAE and diffusion model use stochastic latent representations.

      A published solution is not available for this question yet.

      Question 9 MSQ · 2.0 marks

      Which of the following statements about Denoising Diffusion Implicit Models (DDIMs) are correct? Select all that apply.
      1. The forward process is non-Markovian, unlike the Markovian forward process in DDPM.
      2. DDIM achieves a tighter ELBO than DDPM.
      3. DDIM requires a separate training procedure from DDPM.
      4. When [[IMAGE:173306d6b8cbd418_6_40]] for all [[IMAGE:173306d6b8cbd418_6_41]] , the DDIM sampling process is deterministic.
        Source diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 10 MSQ · 2.0 marks

      Consider two DDPMs at the same timestep [[IMAGE:173306d6b8cbd418_6_42]] . For Model A, [[IMAGE:173306d6b8cbd418_6_43]] , while for Model B, [[IMAGE:173306d6b8cbd418_6_44]] . Recall that [[IMAGE:173306d6b8cbd418_6_45]] Which of the following statements are correct?
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
      1. The sample [[IMAGE:173306d6b8cbd418_6_46]] of Model A is more strongly influenced by the original sample [[IMAGE:173306d6b8cbd418_6_47]] than that of Model B.
        Source diagram or notationSource diagram or notation
      2. The sample [[IMAGE:173306d6b8cbd418_6_48]] of Model B is closer to [[IMAGE:173306d6b8cbd418_6_49]] than that of Model A.
        Source diagram or notationSource diagram or notation
      3. As [[IMAGE:173306d6b8cbd418_6_50]] decreases, the contribution of [[IMAGE:173306d6b8cbd418_6_51]] to [[IMAGE:173306d6b8cbd418_6_52]] increases.
        Source diagram or notationSource diagram or notationSource diagram or notation
      4. Model B produces a noisier representation at timestep [[IMAGE:173306d6b8cbd418_6_53]] than Model A.
        Source diagram or notation

      A published solution is not available for this question yet.

      Question 11 MSQ · 2.0 marks

      Which of the following statements correctly describe the forward diffusion process in a DDPM?
      1. The forward process gradually corrupts [[IMAGE:173306d6b8cbd418_7_54]] by adding Gaussian noise according to a predefined noise schedule.
        Source diagram or notation
      2. The parameters of the forward process are learned using the training data.
      3. The forward process requires a neural network to predict the noise added at every timestep.
      4. Given [[IMAGE:173306d6b8cbd418_7_55]] , [[IMAGE:173306d6b8cbd418_7_56]] can be sampled directly without simulating all intermediate timesteps.
        Source diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 12 MSQ · 2.0 marks

      You are training a GAN with generator [[IMAGE:173306d6b8cbd418_7_57]] , [[IMAGE:173306d6b8cbd418_7_58]] , to generate handwritten digit images spanning all 10 digit classes. After training, you suspect that [[IMAGE:173306d6b8cbd418_7_59]] has not converged to [[IMAGE:173306d6b8cbd418_7_60]] and that the generator is suffering from mode collapse. Which of the following could be indicators of this problem? Select all that apply.
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
      1. Samples [[IMAGE:173306d6b8cbd418_7_61]] for many different [[IMAGE:173306d6b8cbd418_7_62]] all resemble the digit "3".
        Source diagram or notationSource diagram or notation
      2. The discriminator loss [[IMAGE:173306d6b8cbd418_7_63]] drops close to zero while the generator loss [[IMAGE:173306d6b8cbd418_7_64]] remains persistently high.
        Source diagram or notationSource diagram or notation
      3. The discriminator loss remains persistently high while the generator loss drops to zero.
      4. The generator loss and discriminator loss both decrease monotonically to zero, indicating the adversarial objective [[IMAGE:173306d6b8cbd418_7_65]] has reached a global saddle point.
        Source diagram or notation

      A published solution is not available for this question yet.

      Question 13 MSQ · 5.0 marks

      [[IMAGE:173306d6b8cbd418_8_66]]
      Source diagram or notation
      1. [[IMAGE:173306d6b8cbd418_8_67]] : both actions are equally good relative to the policy average in state B, so no policy gradient update will change [[IMAGE:173306d6b8cbd418_8_68]] .
        Source diagram or notationSource diagram or notation
      2. In state C, action [[IMAGE:173306d6b8cbd418_8_69]] has a higher advantage than action [[IMAGE:173306d6b8cbd418_8_70]] ; so a policy gradient update increases [[IMAGE:173306d6b8cbd418_8_71]] and decreases [[IMAGE:173306d6b8cbd418_8_72]] .
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
      3. In state C, action [[IMAGE:173306d6b8cbd418_8_73]] has a higher advantage than action [[IMAGE:173306d6b8cbd418_8_74]] ; so a policy gradient update decreases [[IMAGE:173306d6b8cbd418_8_75]] and increases [[IMAGE:173306d6b8cbd418_8_76]] .
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
      4. If a sampled trajectory from state [[IMAGE:173306d6b8cbd418_8_77]] yields the discounted return [[IMAGE:173306d6b8cbd418_8_78]] , then the estimated advantage [[IMAGE:173306d6b8cbd418_8_79]] ; so the policy gradient update will decrease [[IMAGE:173306d6b8cbd418_8_80]] .
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
      5. If a sampled trajectory from state [[IMAGE:173306d6b8cbd418_8_81]] yields the discounted return [[IMAGE:173306d6b8cbd418_8_82]] , then the estimated advantage [[IMAGE:173306d6b8cbd418_8_83]] ; so the policy gradient update will increase [[IMAGE:173306d6b8cbd418_8_84]] .
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
      6. The expected return obtained by taking [[IMAGE:173306d6b8cbd418_8_85]] in state [[IMAGE:173306d6b8cbd418_8_86]] must be negative.
        Source diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 14 NAT · 3.0 marks

      A transformer processes a sequence of [[IMAGE:173306d6b8cbd418_9_87]] tokens with model dimension [[IMAGE:173306d6b8cbd418_9_88]] using [[IMAGE:173306d6b8cbd418_9_89]] attention heads, each with [[IMAGE:173306d6b8cbd418_9_90]] . The input embedding matrix is: [[IMAGE:173306d6b8cbd418_9_91]] Head 1 uses projections: [[IMAGE:173306d6b8cbd418_9_92]] Head 2 uses projections: [[IMAGE:173306d6b8cbd418_9_93]] Use [[IMAGE:173306d6b8cbd418_9_94]] for scaling. All biases are zero. Based on the above data, answer the given subquestions.
      [[IMAGE:173306d6b8cbd418_9_98]] Compute [[IMAGE:173306d6b8cbd418_9_95]] , [[IMAGE:173306d6b8cbd418_9_96]] , [[IMAGE:173306d6b8cbd418_9_97]] for Head 1 and the scaled score matrix , and enter the value of the [[IMAGE:173306d6b8cbd418_9_99]] entry of [[IMAGE:173306d6b8cbd418_9_100]] . Enter the answer correct to two decimal places.
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 15 NAT · 3.0 marks

        A transformer processes a sequence of [[IMAGE:173306d6b8cbd418_9_87]] tokens with model dimension [[IMAGE:173306d6b8cbd418_9_88]] using [[IMAGE:173306d6b8cbd418_9_89]] attention heads, each with [[IMAGE:173306d6b8cbd418_9_90]] . The input embedding matrix is: [[IMAGE:173306d6b8cbd418_9_91]] Head 1 uses projections: [[IMAGE:173306d6b8cbd418_9_92]] Head 2 uses projections: [[IMAGE:173306d6b8cbd418_9_93]] Use [[IMAGE:173306d6b8cbd418_9_94]] for scaling. All biases are zero. Based on the above data, answer the given subquestions.
        Causal masking is applied on [[IMAGE:173306d6b8cbd418_10_101]] . Compute the attention weights for the third query and enter the value of the largest weight among the three, correct to two decimal places.
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 16 NAT · 3.0 marks

          A transformer processes a sequence of [[IMAGE:173306d6b8cbd418_9_87]] tokens with model dimension [[IMAGE:173306d6b8cbd418_9_88]] using [[IMAGE:173306d6b8cbd418_9_89]] attention heads, each with [[IMAGE:173306d6b8cbd418_9_90]] . The input embedding matrix is: [[IMAGE:173306d6b8cbd418_9_91]] Head 1 uses projections: [[IMAGE:173306d6b8cbd418_9_92]] Head 2 uses projections: [[IMAGE:173306d6b8cbd418_9_93]] Use [[IMAGE:173306d6b8cbd418_9_94]] for scaling. All biases are zero. Based on the above data, answer the given subquestions.
          Using causal masking, compute the attention outputs for the third token from both heads and concatenate them as [[IMAGE:173306d6b8cbd418_10_102]] Enter the fourth entry of the third row of [[IMAGE:173306d6b8cbd418_10_103]] . Enter the answer correct to two decimal places.
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 17 NAT · 3.0 marks

            In classifier-free guidance, a single diffusion model is jointly trained for both conditional and unconditional generation. During training, the conditioning information [[IMAGE:173306d6b8cbd418_11_104]] is randomly replaced with a null token [[IMAGE:173306d6b8cbd418_11_105]] with probability [[IMAGE:173306d6b8cbd418_11_106]] . During sampling at timestep [[IMAGE:173306d6b8cbd418_11_107]] , the model produces two noise predictions: 1. Conditional prediction: [[IMAGE:173306d6b8cbd418_11_108]] (model receives conditioning [[IMAGE:173306d6b8cbd418_11_109]] ) 2. Unconditional prediction: [[IMAGE:173306d6b8cbd418_11_110]] (model receives null token) The classifier-free guided noise prediction is defined as : [[IMAGE:173306d6b8cbd418_11_111]] where [[IMAGE:173306d6b8cbd418_11_112]] is the guidance scale. This follows from the guided score formulation: [[IMAGE:173306d6b8cbd418_11_113]] [[IMAGE:173306d6b8cbd418_11_114]] The predicted clean data point is recovered via . Following are the values given: [[IMAGE:173306d6b8cbd418_11_115]] Based on the above data, answer the given subquestions.
            Compute the predicted clean data point [[IMAGE:173306d6b8cbd418_11_116]] . Enter the answer correct to one decimal place.
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 18 MCQ · 2.0 marks

              In classifier-free guidance, a single diffusion model is jointly trained for both conditional and unconditional generation. During training, the conditioning information [[IMAGE:173306d6b8cbd418_11_104]] is randomly replaced with a null token [[IMAGE:173306d6b8cbd418_11_105]] with probability [[IMAGE:173306d6b8cbd418_11_106]] . During sampling at timestep [[IMAGE:173306d6b8cbd418_11_107]] , the model produces two noise predictions: 1. Conditional prediction: [[IMAGE:173306d6b8cbd418_11_108]] (model receives conditioning [[IMAGE:173306d6b8cbd418_11_109]] ) 2. Unconditional prediction: [[IMAGE:173306d6b8cbd418_11_110]] (model receives null token) The classifier-free guided noise prediction is defined as : [[IMAGE:173306d6b8cbd418_11_111]] where [[IMAGE:173306d6b8cbd418_11_112]] is the guidance scale. This follows from the guided score formulation: [[IMAGE:173306d6b8cbd418_11_113]] [[IMAGE:173306d6b8cbd418_11_114]] The predicted clean data point is recovered via . Following are the values given: [[IMAGE:173306d6b8cbd418_11_115]] Based on the above data, answer the given subquestions.
              If the guidance scale is set to [[IMAGE:173306d6b8cbd418_12_117]] , which of the following correctly describes the effect on [[IMAGE:173306d6b8cbd418_12_118]] and [[IMAGE:173306d6b8cbd418_12_119]] ?
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
              1. [[IMAGE:173306d6b8cbd418_12_120]] reduces to the conditional prediction [[IMAGE:173306d6b8cbd418_12_121]] alone, giving [[IMAGE:173306d6b8cbd418_12_122]] .
                Source diagram or notationSource diagram or notationSource diagram or notation
              2. [[IMAGE:173306d6b8cbd418_12_123]] reduces to the unconditional prediction [[IMAGE:173306d6b8cbd418_12_124]] alone, giving [[IMAGE:173306d6b8cbd418_12_125]] .
                Source diagram or notationSource diagram or notationSource diagram or notation
              3. [[IMAGE:173306d6b8cbd418_12_126]] reduces to the conditional prediction [[IMAGE:173306d6b8cbd418_12_127]] alone, giving [[IMAGE:173306d6b8cbd418_12_128]] .
                Source diagram or notationSource diagram or notationSource diagram or notation
              4. [[IMAGE:173306d6b8cbd418_12_129]] remains unchanged from [[IMAGE:173306d6b8cbd418_12_130]] case because the guidance scale only affects the output projection layer of the U-Net.
                Source diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 19 NAT · 4.0 marks

              Consider a 1-dimensional VAE with: 1. Encoder: [[IMAGE:173306d6b8cbd418_12_131]] , [[IMAGE:173306d6b8cbd418_12_132]] , [[IMAGE:173306d6b8cbd418_12_133]] 2. Decoder: [[IMAGE:173306d6b8cbd418_12_134]] , with fixed decoder variance [[IMAGE:173306d6b8cbd418_12_135]] Use the reparameterization trick to express the latent sample [[IMAGE:173306d6b8cbd418_12_136]] . The reconstruction loss [[IMAGE:173306d6b8cbd418_12_137]] for a single sample is given by [[IMAGE:173306d6b8cbd418_12_138]] . Following are the values given: [[IMAGE:173306d6b8cbd418_12_139]] , input [[IMAGE:173306d6b8cbd418_12_140]] , sampled noise [[IMAGE:173306d6b8cbd418_12_141]] . Compute [[IMAGE:173306d6b8cbd418_12_142]] . Enter the answer correct to three decimal places.
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                A published solution is not available for this question yet.

                Question 20 MCQ · 1.0 marks

                Inference in DDPM is much slower as compared to GAN or VAE.
                1. True
                2. False

                A published solution is not available for this question yet.