MauryaHub PYQ Practice

da5002_2026T1_Q2_NA.pdf

Mathematical Foundations of Generative AI · Quiz 2 · Jan 2026

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MSQ · 3.0 marks

Which among the following is true for the EM algorithm?
  1. At any iteration of the EM algorithm, the value of the objective function never decreases.
  2. It optimizes an upper bound of the objective function.
  3. It optimizes a lower bound of the objective function.
  4. The EM algorithm is guaranteed to give a global solution.
  5. Upon convergence of the EM algorithm for a Gaussian Mixture Model, the mixture proportions [[IMAGE:fa8e6e04a659df48_2_2]] become equal to the posterior probabilities [[IMAGE:fa8e6e04a659df48_2_3]] for all training sample [[IMAGE:fa8e6e04a659df48_2_4]] .
    Source diagram or notationSource diagram or notationSource diagram or notation

A published solution is not available for this question yet.

Question 3 NAT · 1.0 marks

Let [[IMAGE:fa8e6e04a659df48_2_5]] iid [[IMAGE:fa8e6e04a659df48_2_6]] , where [[IMAGE:fa8e6e04a659df48_2_7]] is a 2-component Gaussian mixture distribution with density [[IMAGE:fa8e6e04a659df48_2_8]] where [[IMAGE:fa8e6e04a659df48_2_9]] , and [[IMAGE:fa8e6e04a659df48_2_10]] is known. Based on the above data, answer the given subquestions.
Find the number of parameters to be estimated in the model.
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 4 NAT · 4.0 marks

    Let [[IMAGE:fa8e6e04a659df48_2_5]] iid [[IMAGE:fa8e6e04a659df48_2_6]] , where [[IMAGE:fa8e6e04a659df48_2_7]] is a 2-component Gaussian mixture distribution with density [[IMAGE:fa8e6e04a659df48_2_8]] where [[IMAGE:fa8e6e04a659df48_2_9]] , and [[IMAGE:fa8e6e04a659df48_2_10]] is known. Based on the above data, answer the given subquestions.
    Suppose in the M-step of the EM algorithm, we have the following responsibilities for data points [[IMAGE:fa8e6e04a659df48_3_11]] and [[IMAGE:fa8e6e04a659df48_3_12]] : [[IMAGE:fa8e6e04a659df48_3_13]] [[IMAGE:fa8e6e04a659df48_3_14]] where [[IMAGE:fa8e6e04a659df48_3_15]] Calculate the updated mean [[IMAGE:fa8e6e04a659df48_3_16]] for the second component. Enter the answer correct to two decimal places.
    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 5 NAT · 2.0 marks

      For a latent variable model, the Evidence Lower Bound (ELBO) is given by: [[IMAGE:fa8e6e04a659df48_4_17]] Given a data point [[IMAGE:fa8e6e04a659df48_4_18]] , an approximate posterior [[IMAGE:fa8e6e04a659df48_4_19]] , and the log- joint distribution values [[IMAGE:fa8e6e04a659df48_4_20]] and [[IMAGE:fa8e6e04a659df48_4_21]] , what is the value of the ELBO? Enter the answer correct to two decimal places. Use [[IMAGE:fa8e6e04a659df48_4_22]] for the calculation.
      Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 6 NAT · 2.0 marks

        The DDPM forward process is defined by [[IMAGE:fa8e6e04a659df48_4_23]] , where [[IMAGE:fa8e6e04a659df48_4_24]] and [[IMAGE:fa8e6e04a659df48_4_25]] . For timestep [[IMAGE:fa8e6e04a659df48_4_26]] , [[IMAGE:fa8e6e04a659df48_4_27]] . Given a normalized data point [[IMAGE:fa8e6e04a659df48_4_28]] and a noise sample [[IMAGE:fa8e6e04a659df48_4_29]] , what is the value of [[IMAGE:fa8e6e04a659df48_4_30]] ? Enter the answer correct to two decimal places.
        Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 7 MCQ · 3.0 marks

          A factory uses a Variational Autoencoder (VAE) to automate the inspection of high-precision ceramic tiles. The model is trained to identify defects by learning the distribution of ''normal" tiles. For each input image [[IMAGE:fa8e6e04a659df48_5_31]] (resized to [[IMAGE:fa8e6e04a659df48_5_32]] pixels), the encoder [[IMAGE:fa8e6e04a659df48_5_33]] maps the image to a [[IMAGE:fa8e6e04a659df48_5_34]] dimensional latent space, while the decoder [[IMAGE:fa8e6e04a659df48_5_35]] attempts to reconstruct the original pixels. The training objective to maximize the Evidence Lower Bound (ELBO), defined as: [[IMAGE:fa8e6e04a659df48_5_36]] Assume the prior [[IMAGE:fa8e6e04a659df48_5_37]] is a standard normal distribution [[IMAGE:fa8e6e04a659df48_5_38]] . [[IMAGE:fa8e6e04a659df48_5_39]]
          [[IMAGE:fa8e6e04a659df48_5_40]]
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
          1. [[IMAGE:fa8e6e04a659df48_5_41]]
            Source diagram or notation
          2. [[IMAGE:fa8e6e04a659df48_5_42]]
            Source diagram or notation
          3. [[IMAGE:fa8e6e04a659df48_5_43]]
            Source diagram or notation
          4. [[IMAGE:fa8e6e04a659df48_5_44]]
            Source diagram or notation

          A published solution is not available for this question yet.

          Question 8 NAT · 3.0 marks

          A factory uses a Variational Autoencoder (VAE) to automate the inspection of high-precision ceramic tiles. The model is trained to identify defects by learning the distribution of ''normal" tiles. For each input image [[IMAGE:fa8e6e04a659df48_5_31]] (resized to [[IMAGE:fa8e6e04a659df48_5_32]] pixels), the encoder [[IMAGE:fa8e6e04a659df48_5_33]] maps the image to a [[IMAGE:fa8e6e04a659df48_5_34]] dimensional latent space, while the decoder [[IMAGE:fa8e6e04a659df48_5_35]] attempts to reconstruct the original pixels. The training objective to maximize the Evidence Lower Bound (ELBO), defined as: [[IMAGE:fa8e6e04a659df48_5_36]] Assume the prior [[IMAGE:fa8e6e04a659df48_5_37]] is a standard normal distribution [[IMAGE:fa8e6e04a659df48_5_38]] . [[IMAGE:fa8e6e04a659df48_5_39]]
          Compute the latent regularization contribution [[IMAGE:fa8e6e04a659df48_6_45]] for this tile. Enter the answer correct to two decimal places.
          Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 9 NAT · 2.0 marks

            A factory uses a Variational Autoencoder (VAE) to automate the inspection of high-precision ceramic tiles. The model is trained to identify defects by learning the distribution of ''normal" tiles. For each input image [[IMAGE:fa8e6e04a659df48_5_31]] (resized to [[IMAGE:fa8e6e04a659df48_5_32]] pixels), the encoder [[IMAGE:fa8e6e04a659df48_5_33]] maps the image to a [[IMAGE:fa8e6e04a659df48_5_34]] dimensional latent space, while the decoder [[IMAGE:fa8e6e04a659df48_5_35]] attempts to reconstruct the original pixels. The training objective to maximize the Evidence Lower Bound (ELBO), defined as: [[IMAGE:fa8e6e04a659df48_5_36]] Assume the prior [[IMAGE:fa8e6e04a659df48_5_37]] is a standard normal distribution [[IMAGE:fa8e6e04a659df48_5_38]] . [[IMAGE:fa8e6e04a659df48_5_39]]
            The total reconstruction mismatch ( negative log-likelihood ) for the training sample is found to be [[IMAGE:fa8e6e04a659df48_6_46]] . Calculate the ELBO objective value [[IMAGE:fa8e6e04a659df48_6_47]] for this sample. Round your answer to the nearest integer.
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 10 NAT · 2.0 marks

              A factory uses a Variational Autoencoder (VAE) to automate the inspection of high-precision ceramic tiles. The model is trained to identify defects by learning the distribution of ''normal" tiles. For each input image [[IMAGE:fa8e6e04a659df48_5_31]] (resized to [[IMAGE:fa8e6e04a659df48_5_32]] pixels), the encoder [[IMAGE:fa8e6e04a659df48_5_33]] maps the image to a [[IMAGE:fa8e6e04a659df48_5_34]] dimensional latent space, while the decoder [[IMAGE:fa8e6e04a659df48_5_35]] attempts to reconstruct the original pixels. The training objective to maximize the Evidence Lower Bound (ELBO), defined as: [[IMAGE:fa8e6e04a659df48_5_36]] Assume the prior [[IMAGE:fa8e6e04a659df48_5_37]] is a standard normal distribution [[IMAGE:fa8e6e04a659df48_5_38]] . [[IMAGE:fa8e6e04a659df48_5_39]]
              During a testing phase, the engineers employ a [[IMAGE:fa8e6e04a659df48_6_48]] -VAE variant with [[IMAGE:fa8e6e04a659df48_6_49]] . For a suspicious tile, the reconstruction mismatch is [[IMAGE:fa8e6e04a659df48_6_50]] and the [[IMAGE:fa8e6e04a659df48_6_51]] contribution is [[IMAGE:fa8e6e04a659df48_6_52]] . Compute the [[IMAGE:fa8e6e04a659df48_6_53]] -VAE objective value [[IMAGE:fa8e6e04a659df48_6_54]] for this sample.
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                A published solution is not available for this question yet.

                Question 11 MCQ · 2.0 marks

                A factory uses a Variational Autoencoder (VAE) to automate the inspection of high-precision ceramic tiles. The model is trained to identify defects by learning the distribution of ''normal" tiles. For each input image [[IMAGE:fa8e6e04a659df48_5_31]] (resized to [[IMAGE:fa8e6e04a659df48_5_32]] pixels), the encoder [[IMAGE:fa8e6e04a659df48_5_33]] maps the image to a [[IMAGE:fa8e6e04a659df48_5_34]] dimensional latent space, while the decoder [[IMAGE:fa8e6e04a659df48_5_35]] attempts to reconstruct the original pixels. The training objective to maximize the Evidence Lower Bound (ELBO), defined as: [[IMAGE:fa8e6e04a659df48_5_36]] Assume the prior [[IMAGE:fa8e6e04a659df48_5_37]] is a standard normal distribution [[IMAGE:fa8e6e04a659df48_5_38]] . [[IMAGE:fa8e6e04a659df48_5_39]]
                The factory deployment uses an anomaly score [[IMAGE:fa8e6e04a659df48_7_55]] defined as: [[IMAGE:fa8e6e04a659df48_7_56]] If [[IMAGE:fa8e6e04a659df48_7_57]] , the tile is flagged. Determine the anomaly score for the suspicious tile (reconstruction mismatch: [[IMAGE:fa8e6e04a659df48_7_58]] , [[IMAGE:fa8e6e04a659df48_7_59]] ) and state whether it is flagged or not.
                Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
                1. Flagged
                2. Not flagged

                A published solution is not available for this question yet.

                Question 12 MCQ · 2.0 marks

                A VQ-VAE uses a codebook with [[IMAGE:fa8e6e04a659df48_7_60]] vectors, each of dimension [[IMAGE:fa8e6e04a659df48_7_61]] . The encoder output [[IMAGE:fa8e6e04a659df48_7_62]] is a tensor of shape [[IMAGE:fa8e6e04a659df48_7_63]] . What is the total size (in bits) of the discrete latent representation (the indices sent to the decoder) for a single input?
                Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
                1. 1024 bits
                2. 2048 bits
                3. 4096 bits
                4. 8192 bits

                A published solution is not available for this question yet.

                Question 13 MCQ · 2.0 marks

                In a PyTorch DDPM sampling loop, you have a tensor [[IMAGE:fa8e6e04a659df48_7_64]] of shape [[IMAGE:fa8e6e04a659df48_7_65]] representing the noisy image at step 't'. The model [[IMAGE:fa8e6e04a659df48_7_66]] returns a predicted noise tensor of the same shape. The timestep 't' is a single integer. When passing 't' to the model, it is typically converted to a tensor. What must be the shape of this timestep tensor 't' so that it can be correctly processed by the model for a single image?
                Source diagram or notationSource diagram or notationSource diagram or notation
                1. '[1]'
                2. '[1, 1]'
                3. '[1, 64]'
                4. '[64, 64]'

                A published solution is not available for this question yet.

                Question 14 NAT · 3.0 marks

                In a PyTorch implementation of a VAE encoder for [[IMAGE:fa8e6e04a659df48_8_67]] images, the input is first flattened and then passed through a linear layer 'nn.Linear(3072, 400)'. This is followed by a ReLU activation and then two parallel linear layers to produce [[IMAGE:fa8e6e04a659df48_8_68]] and [[IMAGE:fa8e6e04a659df48_8_69]] , both mapping from 400 features to a latent dimension of 20. What is the total number of trainable weight and bias parameters in the layer that produces the mean vector [[IMAGE:fa8e6e04a659df48_8_70]] ?
                Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                  A published solution is not available for this question yet.

                  Question 15 NAT · 3.0 marks

                  During the inference phase of a DDPM, a trained U-Net [[IMAGE:fa8e6e04a659df48_8_71]] is used to predict the noise added to a sample. Once the noise is predicted, the original data point can be estimated using the formula: [[IMAGE:fa8e6e04a659df48_8_72]] At timestep [[IMAGE:fa8e6e04a659df48_8_73]] , given [[IMAGE:fa8e6e04a659df48_8_74]] and a noisy latent [[IMAGE:fa8e6e04a659df48_8_75]] , the U-Net predicts the noise [[IMAGE:fa8e6e04a659df48_9_76]] . Calculate the estimated clean data point [[IMAGE:fa8e6e04a659df48_9_77]] .
                  Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                    A published solution is not available for this question yet.

                    Question 16 MCQ · 3.0 marks

                    In a trained Denoising Diffusion Probabilistic Model (DDPM), which of the following statements correctly describes the inference (sampling) process to generate a new data point [[IMAGE:fa8e6e04a659df48_9_78]] ?
                    Source diagram or notation
                    1. A new sample is generated in a single forward pass by feeding a latent vector [[IMAGE:fa8e6e04a659df48_9_79]] into the trained U-Net.
                      Source diagram or notation
                    2. To obtain a sample [[IMAGE:fa8e6e04a659df48_9_80]] , the model must iteratively traverse the reverse decoding process [[IMAGE:fa8e6e04a659df48_9_81]] , starting from Gaussian noise [[IMAGE:fa8e6e04a659df48_9_82]] .
                      Source diagram or notationSource diagram or notationSource diagram or notation
                    3. The inference process is faster than training because it does not require any forward passes through the neural network.
                    4. During inference, the U-Net predicts the original data point [[IMAGE:fa8e6e04a659df48_9_83]] directly from [[IMAGE:fa8e6e04a659df48_9_84]] without any intermediate steps.
                      Source diagram or notationSource diagram or notation

                    A published solution is not available for this question yet.

                    Question 17 MCQ · 3.0 marks

                    [[IMAGE:fa8e6e04a659df48_10_85]]
                    Source diagram or notation
                    1. [[IMAGE:fa8e6e04a659df48_10_86]]
                      Source diagram or notation
                    2. [[IMAGE:fa8e6e04a659df48_10_87]]
                      Source diagram or notation
                    3. [[IMAGE:fa8e6e04a659df48_10_88]]
                      Source diagram or notation
                    4. [[IMAGE:fa8e6e04a659df48_10_89]]
                      Source diagram or notation

                    A published solution is not available for this question yet.