da5002_2026T1_Q2_NA.pdf
Mathematical Foundations of Generative AI · Quiz 2 · Jan 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MSQ · 3.0 marks
Which among the following is true for the EM algorithm?
At any iteration of the EM algorithm, the value of the objective function never
decreases.
It optimizes an upper bound of the objective function.
It optimizes a lower bound of the objective function.
The EM algorithm is guaranteed to give a global solution.
Upon convergence of the EM algorithm for a Gaussian Mixture Model, the
mixture proportions [[IMAGE:fa8e6e04a659df48_2_2]] become equal to the posterior probabilities [[IMAGE:fa8e6e04a659df48_2_3]] for all training
sample [[IMAGE:fa8e6e04a659df48_2_4]] .



A published solution is not available for this question yet.
Question 3 NAT · 1.0 marks
Let [[IMAGE:fa8e6e04a659df48_2_5]] iid [[IMAGE:fa8e6e04a659df48_2_6]] , where [[IMAGE:fa8e6e04a659df48_2_7]] is a 2-component Gaussian mixture distribution with density
[[IMAGE:fa8e6e04a659df48_2_8]]
where [[IMAGE:fa8e6e04a659df48_2_9]] , and [[IMAGE:fa8e6e04a659df48_2_10]] is known.
Based on the above data, answer the given subquestions.
Find the number of parameters to be estimated in the model.






A published solution is not available for this question yet.
Question 4 NAT · 4.0 marks
Let [[IMAGE:fa8e6e04a659df48_2_5]] iid [[IMAGE:fa8e6e04a659df48_2_6]] , where [[IMAGE:fa8e6e04a659df48_2_7]] is a 2-component Gaussian mixture distribution with density
[[IMAGE:fa8e6e04a659df48_2_8]]
where [[IMAGE:fa8e6e04a659df48_2_9]] , and [[IMAGE:fa8e6e04a659df48_2_10]] is known.
Based on the above data, answer the given subquestions.
Suppose in the M-step of the EM algorithm, we have the following responsibilities for data points
[[IMAGE:fa8e6e04a659df48_3_11]] and [[IMAGE:fa8e6e04a659df48_3_12]] :
[[IMAGE:fa8e6e04a659df48_3_13]]
[[IMAGE:fa8e6e04a659df48_3_14]]
where [[IMAGE:fa8e6e04a659df48_3_15]] Calculate the updated mean [[IMAGE:fa8e6e04a659df48_3_16]] for the second component. Enter the
answer correct to two decimal places.












A published solution is not available for this question yet.
Question 5 NAT · 2.0 marks
For a latent variable model, the Evidence Lower Bound (ELBO) is given by:
[[IMAGE:fa8e6e04a659df48_4_17]]
Given a data point [[IMAGE:fa8e6e04a659df48_4_18]] , an approximate posterior [[IMAGE:fa8e6e04a659df48_4_19]] , and the log-
joint distribution values [[IMAGE:fa8e6e04a659df48_4_20]] and [[IMAGE:fa8e6e04a659df48_4_21]] , what is the value of
the ELBO? Enter the answer correct to two decimal places. Use [[IMAGE:fa8e6e04a659df48_4_22]] for the calculation.






A published solution is not available for this question yet.
Question 6 NAT · 2.0 marks
The DDPM forward process is defined by [[IMAGE:fa8e6e04a659df48_4_23]] , where [[IMAGE:fa8e6e04a659df48_4_24]] and
[[IMAGE:fa8e6e04a659df48_4_25]]
. For timestep [[IMAGE:fa8e6e04a659df48_4_26]] , [[IMAGE:fa8e6e04a659df48_4_27]] . Given a normalized data point [[IMAGE:fa8e6e04a659df48_4_28]] and a
noise sample [[IMAGE:fa8e6e04a659df48_4_29]] , what is the value of [[IMAGE:fa8e6e04a659df48_4_30]] ? Enter the answer correct to two decimal places.








A published solution is not available for this question yet.
Question 7 MCQ · 3.0 marks
A factory uses a Variational Autoencoder (VAE) to automate the inspection of high-precision
ceramic tiles. The model is trained to identify defects by learning the distribution of ''normal" tiles.
For each input image [[IMAGE:fa8e6e04a659df48_5_31]] (resized to [[IMAGE:fa8e6e04a659df48_5_32]] pixels), the encoder [[IMAGE:fa8e6e04a659df48_5_33]] maps the image to a
[[IMAGE:fa8e6e04a659df48_5_34]] dimensional latent space, while the decoder [[IMAGE:fa8e6e04a659df48_5_35]] attempts to reconstruct the original
pixels. The training objective to maximize the Evidence Lower Bound (ELBO), defined as:
[[IMAGE:fa8e6e04a659df48_5_36]]
Assume the prior [[IMAGE:fa8e6e04a659df48_5_37]] is a standard normal distribution [[IMAGE:fa8e6e04a659df48_5_38]] .
[[IMAGE:fa8e6e04a659df48_5_39]]
[[IMAGE:fa8e6e04a659df48_5_40]]










[[IMAGE:fa8e6e04a659df48_5_41]]

[[IMAGE:fa8e6e04a659df48_5_42]]

[[IMAGE:fa8e6e04a659df48_5_43]]

[[IMAGE:fa8e6e04a659df48_5_44]]

A published solution is not available for this question yet.
Question 8 NAT · 3.0 marks
A factory uses a Variational Autoencoder (VAE) to automate the inspection of high-precision
ceramic tiles. The model is trained to identify defects by learning the distribution of ''normal" tiles.
For each input image [[IMAGE:fa8e6e04a659df48_5_31]] (resized to [[IMAGE:fa8e6e04a659df48_5_32]] pixels), the encoder [[IMAGE:fa8e6e04a659df48_5_33]] maps the image to a
[[IMAGE:fa8e6e04a659df48_5_34]] dimensional latent space, while the decoder [[IMAGE:fa8e6e04a659df48_5_35]] attempts to reconstruct the original
pixels. The training objective to maximize the Evidence Lower Bound (ELBO), defined as:
[[IMAGE:fa8e6e04a659df48_5_36]]
Assume the prior [[IMAGE:fa8e6e04a659df48_5_37]] is a standard normal distribution [[IMAGE:fa8e6e04a659df48_5_38]] .
[[IMAGE:fa8e6e04a659df48_5_39]]
Compute the latent regularization contribution [[IMAGE:fa8e6e04a659df48_6_45]] for this tile. Enter the answer correct to two
decimal places.










A published solution is not available for this question yet.
Question 9 NAT · 2.0 marks
A factory uses a Variational Autoencoder (VAE) to automate the inspection of high-precision
ceramic tiles. The model is trained to identify defects by learning the distribution of ''normal" tiles.
For each input image [[IMAGE:fa8e6e04a659df48_5_31]] (resized to [[IMAGE:fa8e6e04a659df48_5_32]] pixels), the encoder [[IMAGE:fa8e6e04a659df48_5_33]] maps the image to a
[[IMAGE:fa8e6e04a659df48_5_34]] dimensional latent space, while the decoder [[IMAGE:fa8e6e04a659df48_5_35]] attempts to reconstruct the original
pixels. The training objective to maximize the Evidence Lower Bound (ELBO), defined as:
[[IMAGE:fa8e6e04a659df48_5_36]]
Assume the prior [[IMAGE:fa8e6e04a659df48_5_37]] is a standard normal distribution [[IMAGE:fa8e6e04a659df48_5_38]] .
[[IMAGE:fa8e6e04a659df48_5_39]]
The total reconstruction mismatch ( negative log-likelihood ) for the training sample is found to be
[[IMAGE:fa8e6e04a659df48_6_46]] . Calculate the ELBO objective value [[IMAGE:fa8e6e04a659df48_6_47]] for this sample. Round your answer to the
nearest integer.











A published solution is not available for this question yet.
Question 10 NAT · 2.0 marks
A factory uses a Variational Autoencoder (VAE) to automate the inspection of high-precision
ceramic tiles. The model is trained to identify defects by learning the distribution of ''normal" tiles.
For each input image [[IMAGE:fa8e6e04a659df48_5_31]] (resized to [[IMAGE:fa8e6e04a659df48_5_32]] pixels), the encoder [[IMAGE:fa8e6e04a659df48_5_33]] maps the image to a
[[IMAGE:fa8e6e04a659df48_5_34]] dimensional latent space, while the decoder [[IMAGE:fa8e6e04a659df48_5_35]] attempts to reconstruct the original
pixels. The training objective to maximize the Evidence Lower Bound (ELBO), defined as:
[[IMAGE:fa8e6e04a659df48_5_36]]
Assume the prior [[IMAGE:fa8e6e04a659df48_5_37]] is a standard normal distribution [[IMAGE:fa8e6e04a659df48_5_38]] .
[[IMAGE:fa8e6e04a659df48_5_39]]
During a testing phase, the engineers employ a [[IMAGE:fa8e6e04a659df48_6_48]] -VAE variant with [[IMAGE:fa8e6e04a659df48_6_49]] . For a suspicious tile,
the reconstruction mismatch is [[IMAGE:fa8e6e04a659df48_6_50]] and the [[IMAGE:fa8e6e04a659df48_6_51]] contribution is [[IMAGE:fa8e6e04a659df48_6_52]] . Compute the [[IMAGE:fa8e6e04a659df48_6_53]] -VAE
objective value [[IMAGE:fa8e6e04a659df48_6_54]] for this sample.
















A published solution is not available for this question yet.
Question 11 MCQ · 2.0 marks
A factory uses a Variational Autoencoder (VAE) to automate the inspection of high-precision
ceramic tiles. The model is trained to identify defects by learning the distribution of ''normal" tiles.
For each input image [[IMAGE:fa8e6e04a659df48_5_31]] (resized to [[IMAGE:fa8e6e04a659df48_5_32]] pixels), the encoder [[IMAGE:fa8e6e04a659df48_5_33]] maps the image to a
[[IMAGE:fa8e6e04a659df48_5_34]] dimensional latent space, while the decoder [[IMAGE:fa8e6e04a659df48_5_35]] attempts to reconstruct the original
pixels. The training objective to maximize the Evidence Lower Bound (ELBO), defined as:
[[IMAGE:fa8e6e04a659df48_5_36]]
Assume the prior [[IMAGE:fa8e6e04a659df48_5_37]] is a standard normal distribution [[IMAGE:fa8e6e04a659df48_5_38]] .
[[IMAGE:fa8e6e04a659df48_5_39]]
The factory deployment uses an anomaly score [[IMAGE:fa8e6e04a659df48_7_55]] defined as:
[[IMAGE:fa8e6e04a659df48_7_56]]
If [[IMAGE:fa8e6e04a659df48_7_57]] , the tile is flagged. Determine the anomaly score for the suspicious tile
(reconstruction mismatch: [[IMAGE:fa8e6e04a659df48_7_58]] , [[IMAGE:fa8e6e04a659df48_7_59]] ) and state whether it is flagged or not.














Flagged
Not flagged
A published solution is not available for this question yet.
Question 12 MCQ · 2.0 marks
A VQ-VAE uses a codebook with [[IMAGE:fa8e6e04a659df48_7_60]] vectors, each of dimension [[IMAGE:fa8e6e04a659df48_7_61]] . The encoder
output [[IMAGE:fa8e6e04a659df48_7_62]] is a tensor of shape [[IMAGE:fa8e6e04a659df48_7_63]] . What is the total size (in bits) of the discrete latent
representation (the indices sent to the decoder) for a single input?




1024 bits
2048 bits
4096 bits
8192 bits
A published solution is not available for this question yet.
Question 13 MCQ · 2.0 marks
In a PyTorch DDPM sampling loop, you have a tensor [[IMAGE:fa8e6e04a659df48_7_64]] of shape [[IMAGE:fa8e6e04a659df48_7_65]] representing the
noisy image at step 't'. The model [[IMAGE:fa8e6e04a659df48_7_66]] returns a predicted noise tensor of the same shape.
The timestep 't' is a single integer. When passing 't' to the model, it is typically converted to a
tensor. What must be the shape of this timestep tensor 't' so that it can be correctly processed by
the model for a single image?



'[1]'
'[1, 1]'
'[1, 64]'
'[64, 64]'
A published solution is not available for this question yet.
Question 14 NAT · 3.0 marks
In a PyTorch implementation of a VAE encoder for [[IMAGE:fa8e6e04a659df48_8_67]] images, the input is first flattened
and then passed through a linear layer 'nn.Linear(3072, 400)'. This is followed by a ReLU activation
and then two parallel linear layers to produce [[IMAGE:fa8e6e04a659df48_8_68]] and [[IMAGE:fa8e6e04a659df48_8_69]] , both mapping from 400 features to a
latent dimension of 20. What is the total number of trainable weight and bias parameters in the
layer that produces the mean vector [[IMAGE:fa8e6e04a659df48_8_70]] ?




A published solution is not available for this question yet.
Question 15 NAT · 3.0 marks
During the inference phase of a DDPM, a trained U-Net [[IMAGE:fa8e6e04a659df48_8_71]] is used to predict the noise added
to a sample. Once the noise is predicted, the original data point can be estimated using the
formula:
[[IMAGE:fa8e6e04a659df48_8_72]]
At timestep [[IMAGE:fa8e6e04a659df48_8_73]] , given [[IMAGE:fa8e6e04a659df48_8_74]] and a noisy latent [[IMAGE:fa8e6e04a659df48_8_75]] , the U-Net predicts the noise
[[IMAGE:fa8e6e04a659df48_9_76]] . Calculate the estimated clean data point [[IMAGE:fa8e6e04a659df48_9_77]] .







A published solution is not available for this question yet.
Question 16 MCQ · 3.0 marks
In a trained Denoising Diffusion Probabilistic Model (DDPM), which of the following statements
correctly describes the inference (sampling) process to generate a new data point [[IMAGE:fa8e6e04a659df48_9_78]] ?

A new sample is generated in a single forward pass by feeding a latent vector
[[IMAGE:fa8e6e04a659df48_9_79]] into the trained U-Net.

To obtain a sample [[IMAGE:fa8e6e04a659df48_9_80]] , the model must iteratively traverse the reverse
decoding process [[IMAGE:fa8e6e04a659df48_9_81]] , starting from Gaussian noise [[IMAGE:fa8e6e04a659df48_9_82]] .



The inference process is faster than training because it does not require any
forward passes through the neural network.
During inference, the U-Net predicts the original data point [[IMAGE:fa8e6e04a659df48_9_83]] directly from
[[IMAGE:fa8e6e04a659df48_9_84]] without any intermediate steps.


A published solution is not available for this question yet.
Question 17 MCQ · 3.0 marks
[[IMAGE:fa8e6e04a659df48_10_85]]

[[IMAGE:fa8e6e04a659df48_10_86]]

[[IMAGE:fa8e6e04a659df48_10_87]]

[[IMAGE:fa8e6e04a659df48_10_88]]

[[IMAGE:fa8e6e04a659df48_10_89]]

A published solution is not available for this question yet.