da5002_2026T2_ET_FN.pdf
Mathematical Foundations of Generative AI · End Term · May 2026 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 NAT · 2.0 marks
An agent receives rewards [[IMAGE:173306d6b8cbd418_2_2]] over an episode ( [[IMAGE:173306d6b8cbd418_2_3]] ) with
the discount factor equal to 0.5. Compute the full discounted return [[IMAGE:173306d6b8cbd418_2_4]] for the trajectory [[IMAGE:173306d6b8cbd418_2_5]] . Enter
your answer correct to three decimal places.




A published solution is not available for this question yet.
Question 3 MSQ · 3.0 marks
[[IMAGE:173306d6b8cbd418_3_6]]

Both TRPO and PPO aim to prevent the updated policy from deviating too far
from [[IMAGE:173306d6b8cbd418_3_7]] .

PPO always produces a smaller policy update than TRPO for the same [[IMAGE:173306d6b8cbd418_3_8]] and [[IMAGE:173306d6b8cbd418_3_9]] .


When [[IMAGE:173306d6b8cbd418_3_10]] is within [[IMAGE:173306d6b8cbd418_3_11]] , the PPO clip is inactive and [[IMAGE:173306d6b8cbd418_3_12]] reduces to
the unclipped surrogate [[IMAGE:173306d6b8cbd418_3_13]] .




Computing the PPO objective requires solving a constrained optimisation
subproblem at each step, making it more expensive than TRPO.
A published solution is not available for this question yet.
Question 4 MSQ · 3.0 marks
In Proximal Policy Optimization (PPO), the clipped surrogate objective for a single state-action pair
is:
[[IMAGE:173306d6b8cbd418_3_14]]
Which of the following statements are correct?

When [[IMAGE:173306d6b8cbd418_3_15]] and [[IMAGE:173306d6b8cbd418_3_16]] , the clipped objective prevents the gradient
from further increasing [[IMAGE:173306d6b8cbd418_3_17]] , effectively imposing a trust region.



When [[IMAGE:173306d6b8cbd418_4_18]] and [[IMAGE:173306d6b8cbd418_4_19]] , the clipped objective prevents the gradient
from further decreasing [[IMAGE:173306d6b8cbd418_4_20]] .



The clipping mechanism is equivalent to adding an explicit KL-divergence
penalty [[IMAGE:173306d6b8cbd418_4_21]] to the objective.

If [[IMAGE:173306d6b8cbd418_4_22]] , then [[IMAGE:173306d6b8cbd418_4_23]] regardless of the sign of [[IMAGE:173306d6b8cbd418_4_24]] .



A published solution is not available for this question yet.
Question 5 MSQ · 3.0 marks
In multi-head self-attention with [[IMAGE:173306d6b8cbd418_4_25]] heads, the queries, keys, and values for each head have
dimension [[IMAGE:173306d6b8cbd418_4_26]] where [[IMAGE:173306d6b8cbd418_4_27]] is the model dimension. Which of the following
correctly explain why this choice makes the implementation efficient?



Choosing [[IMAGE:173306d6b8cbd418_4_28]] ensures that every head learns the same attention
pattern while reducing the number of parameters.

The total parameters across [[IMAGE:173306d6b8cbd418_4_29]] query projections [[IMAGE:173306d6b8cbd418_4_30]] equal [[IMAGE:173306d6b8cbd418_4_31]] ,
the same as a single-head projection [[IMAGE:173306d6b8cbd418_4_32]] . The same holds for key and value
projections.




Since [[IMAGE:173306d6b8cbd418_4_33]] , all [[IMAGE:173306d6b8cbd418_4_34]] heads can be computed as a single batched matrix
multiplication, enabling parallel execution across heads on a GPU.


Using [[IMAGE:173306d6b8cbd418_4_35]] , the total computational cost across all [[IMAGE:173306d6b8cbd418_4_36]] heads is reduced
by a factor of [[IMAGE:173306d6b8cbd418_4_37]] compared to single-head attention with [[IMAGE:173306d6b8cbd418_4_38]] .




A published solution is not available for this question yet.
Question 6 NAT · 3.0 marks
[[IMAGE:173306d6b8cbd418_5_39]]

A published solution is not available for this question yet.
Question 7 MCQ · 2.0 marks
What is the role of layer normalization in a transformer block?
It eliminates the need for residual connections.
It normalizes the hidden representation across its features for each token,
helping maintain a stable scale of activations during training.
It normalizes the hidden representations across different examples in a mini-
batch.
It introduces the nonlinearity needed by the transformer through a learnable
activation function.
A published solution is not available for this question yet.
Question 8 MCQ · 2.0 marks
Which among the following is true for a VAE and a diffusion model?
Both models use an encoder and decoder to map between the input and a
low-dimensional representation.
Both the encoder and decoder are learned in VAE and diffusion.
The hidden variables in a diffusion model are deterministic, whereas the latent
variable in a VAE is stochastic.
Both VAE and diffusion model use stochastic latent representations.
A published solution is not available for this question yet.
Question 9 MSQ · 2.0 marks
Which of the following statements about Denoising Diffusion Implicit Models (DDIMs) are correct?
Select all that apply.
The forward process is non-Markovian, unlike the Markovian forward process
in DDPM.
DDIM achieves a tighter ELBO than DDPM.
DDIM requires a separate training procedure from DDPM.
When [[IMAGE:173306d6b8cbd418_6_40]] for all [[IMAGE:173306d6b8cbd418_6_41]] , the DDIM sampling process is deterministic.


A published solution is not available for this question yet.
Question 10 MSQ · 2.0 marks
Consider two DDPMs at the same timestep [[IMAGE:173306d6b8cbd418_6_42]] . For Model A, [[IMAGE:173306d6b8cbd418_6_43]] , while for Model B, [[IMAGE:173306d6b8cbd418_6_44]] .
Recall that
[[IMAGE:173306d6b8cbd418_6_45]]
Which of the following statements are correct?




The sample [[IMAGE:173306d6b8cbd418_6_46]] of Model A is more strongly influenced by the original sample
[[IMAGE:173306d6b8cbd418_6_47]] than that of Model B.


The sample [[IMAGE:173306d6b8cbd418_6_48]] of Model B is closer to [[IMAGE:173306d6b8cbd418_6_49]] than that of Model A.


As [[IMAGE:173306d6b8cbd418_6_50]] decreases, the contribution of [[IMAGE:173306d6b8cbd418_6_51]] to [[IMAGE:173306d6b8cbd418_6_52]] increases.



Model B produces a noisier representation at timestep [[IMAGE:173306d6b8cbd418_6_53]] than Model A.

A published solution is not available for this question yet.
Question 11 MSQ · 2.0 marks
Which of the following statements correctly describe the forward diffusion process in a DDPM?
The forward process gradually corrupts [[IMAGE:173306d6b8cbd418_7_54]] by adding Gaussian noise
according to a predefined noise schedule.

The parameters of the forward process are learned using the training data.
The forward process requires a neural network to predict the noise added at
every timestep.
Given [[IMAGE:173306d6b8cbd418_7_55]] , [[IMAGE:173306d6b8cbd418_7_56]] can be sampled directly without simulating all intermediate
timesteps.


A published solution is not available for this question yet.
Question 12 MSQ · 2.0 marks
You are training a GAN with generator [[IMAGE:173306d6b8cbd418_7_57]] , [[IMAGE:173306d6b8cbd418_7_58]] , to generate handwritten digit images
spanning all 10 digit classes. After training, you suspect that [[IMAGE:173306d6b8cbd418_7_59]] has not converged to [[IMAGE:173306d6b8cbd418_7_60]] and that
the generator is suffering from mode collapse. Which of the following could be indicators of this
problem? Select all that apply.




Samples [[IMAGE:173306d6b8cbd418_7_61]] for many different [[IMAGE:173306d6b8cbd418_7_62]] all resemble the digit "3".


The discriminator loss [[IMAGE:173306d6b8cbd418_7_63]] drops close to zero while the generator loss
[[IMAGE:173306d6b8cbd418_7_64]] remains persistently high.


The discriminator loss remains persistently high while the generator loss
drops to zero.
The generator loss and discriminator loss both decrease monotonically to
zero, indicating the adversarial objective [[IMAGE:173306d6b8cbd418_7_65]] has reached a global saddle point.

A published solution is not available for this question yet.
Question 13 MSQ · 5.0 marks
[[IMAGE:173306d6b8cbd418_8_66]]

[[IMAGE:173306d6b8cbd418_8_67]] : both actions are equally good relative to the
policy average in state B, so no policy gradient update will change [[IMAGE:173306d6b8cbd418_8_68]] .


In state C, action [[IMAGE:173306d6b8cbd418_8_69]] has a higher advantage than action [[IMAGE:173306d6b8cbd418_8_70]] ; so a policy
gradient update increases [[IMAGE:173306d6b8cbd418_8_71]] and decreases [[IMAGE:173306d6b8cbd418_8_72]] .




In state C, action [[IMAGE:173306d6b8cbd418_8_73]] has a higher advantage than action [[IMAGE:173306d6b8cbd418_8_74]] ; so a policy
gradient update decreases [[IMAGE:173306d6b8cbd418_8_75]] and increases [[IMAGE:173306d6b8cbd418_8_76]] .




If a sampled trajectory from state [[IMAGE:173306d6b8cbd418_8_77]] yields the discounted return [[IMAGE:173306d6b8cbd418_8_78]] ,
then the estimated advantage [[IMAGE:173306d6b8cbd418_8_79]] ; so the policy gradient update will decrease [[IMAGE:173306d6b8cbd418_8_80]] .




If a sampled trajectory from state [[IMAGE:173306d6b8cbd418_8_81]] yields the discounted return [[IMAGE:173306d6b8cbd418_8_82]] ,
then the estimated advantage [[IMAGE:173306d6b8cbd418_8_83]] ; so the policy gradient update will increase [[IMAGE:173306d6b8cbd418_8_84]] .




The expected return obtained by taking [[IMAGE:173306d6b8cbd418_8_85]] in state [[IMAGE:173306d6b8cbd418_8_86]] must be negative.


A published solution is not available for this question yet.
Question 14 NAT · 3.0 marks
A transformer processes a sequence of [[IMAGE:173306d6b8cbd418_9_87]] tokens with model dimension [[IMAGE:173306d6b8cbd418_9_88]] using [[IMAGE:173306d6b8cbd418_9_89]]
attention heads, each with [[IMAGE:173306d6b8cbd418_9_90]] .
The input embedding matrix is:
[[IMAGE:173306d6b8cbd418_9_91]]
Head 1 uses projections:
[[IMAGE:173306d6b8cbd418_9_92]]
Head 2 uses projections:
[[IMAGE:173306d6b8cbd418_9_93]]
Use [[IMAGE:173306d6b8cbd418_9_94]] for scaling. All biases are zero.
Based on the above data, answer the given subquestions.
[[IMAGE:173306d6b8cbd418_9_98]]
Compute [[IMAGE:173306d6b8cbd418_9_95]] , [[IMAGE:173306d6b8cbd418_9_96]] , [[IMAGE:173306d6b8cbd418_9_97]] for Head 1 and the scaled score matrix , and enter the value of
the [[IMAGE:173306d6b8cbd418_9_99]] entry of [[IMAGE:173306d6b8cbd418_9_100]] . Enter the answer correct to two decimal places.














A published solution is not available for this question yet.
Question 15 NAT · 3.0 marks
A transformer processes a sequence of [[IMAGE:173306d6b8cbd418_9_87]] tokens with model dimension [[IMAGE:173306d6b8cbd418_9_88]] using [[IMAGE:173306d6b8cbd418_9_89]]
attention heads, each with [[IMAGE:173306d6b8cbd418_9_90]] .
The input embedding matrix is:
[[IMAGE:173306d6b8cbd418_9_91]]
Head 1 uses projections:
[[IMAGE:173306d6b8cbd418_9_92]]
Head 2 uses projections:
[[IMAGE:173306d6b8cbd418_9_93]]
Use [[IMAGE:173306d6b8cbd418_9_94]] for scaling. All biases are zero.
Based on the above data, answer the given subquestions.
Causal masking is applied on [[IMAGE:173306d6b8cbd418_10_101]] . Compute the attention weights for the third query and enter the
value of the largest weight among the three, correct to two decimal places.









A published solution is not available for this question yet.
Question 16 NAT · 3.0 marks
A transformer processes a sequence of [[IMAGE:173306d6b8cbd418_9_87]] tokens with model dimension [[IMAGE:173306d6b8cbd418_9_88]] using [[IMAGE:173306d6b8cbd418_9_89]]
attention heads, each with [[IMAGE:173306d6b8cbd418_9_90]] .
The input embedding matrix is:
[[IMAGE:173306d6b8cbd418_9_91]]
Head 1 uses projections:
[[IMAGE:173306d6b8cbd418_9_92]]
Head 2 uses projections:
[[IMAGE:173306d6b8cbd418_9_93]]
Use [[IMAGE:173306d6b8cbd418_9_94]] for scaling. All biases are zero.
Based on the above data, answer the given subquestions.
Using causal masking, compute the attention outputs for the third token from both heads and
concatenate them as
[[IMAGE:173306d6b8cbd418_10_102]]
Enter the fourth entry of the third row of [[IMAGE:173306d6b8cbd418_10_103]] . Enter the answer correct to two decimal places.










A published solution is not available for this question yet.
Question 17 NAT · 3.0 marks
In classifier-free guidance, a single diffusion model is jointly trained for both conditional and
unconditional generation. During training, the conditioning information [[IMAGE:173306d6b8cbd418_11_104]] is randomly replaced
with a null token [[IMAGE:173306d6b8cbd418_11_105]] with probability [[IMAGE:173306d6b8cbd418_11_106]] . During sampling at timestep [[IMAGE:173306d6b8cbd418_11_107]] , the model produces
two noise predictions:
1. Conditional prediction: [[IMAGE:173306d6b8cbd418_11_108]] (model receives conditioning [[IMAGE:173306d6b8cbd418_11_109]] )
2. Unconditional prediction: [[IMAGE:173306d6b8cbd418_11_110]] (model receives null token)
The classifier-free guided noise prediction is defined as :
[[IMAGE:173306d6b8cbd418_11_111]]
where [[IMAGE:173306d6b8cbd418_11_112]] is the guidance scale. This follows from the guided score formulation:
[[IMAGE:173306d6b8cbd418_11_113]]
[[IMAGE:173306d6b8cbd418_11_114]]
The predicted clean data point is recovered via .
Following are the values given:
[[IMAGE:173306d6b8cbd418_11_115]]
Based on the above data, answer the given subquestions.
Compute the predicted clean data point [[IMAGE:173306d6b8cbd418_11_116]] . Enter the answer correct to one decimal place.













A published solution is not available for this question yet.
Question 18 MCQ · 2.0 marks
In classifier-free guidance, a single diffusion model is jointly trained for both conditional and
unconditional generation. During training, the conditioning information [[IMAGE:173306d6b8cbd418_11_104]] is randomly replaced
with a null token [[IMAGE:173306d6b8cbd418_11_105]] with probability [[IMAGE:173306d6b8cbd418_11_106]] . During sampling at timestep [[IMAGE:173306d6b8cbd418_11_107]] , the model produces
two noise predictions:
1. Conditional prediction: [[IMAGE:173306d6b8cbd418_11_108]] (model receives conditioning [[IMAGE:173306d6b8cbd418_11_109]] )
2. Unconditional prediction: [[IMAGE:173306d6b8cbd418_11_110]] (model receives null token)
The classifier-free guided noise prediction is defined as :
[[IMAGE:173306d6b8cbd418_11_111]]
where [[IMAGE:173306d6b8cbd418_11_112]] is the guidance scale. This follows from the guided score formulation:
[[IMAGE:173306d6b8cbd418_11_113]]
[[IMAGE:173306d6b8cbd418_11_114]]
The predicted clean data point is recovered via .
Following are the values given:
[[IMAGE:173306d6b8cbd418_11_115]]
Based on the above data, answer the given subquestions.
If the guidance scale is set to [[IMAGE:173306d6b8cbd418_12_117]] , which of the following correctly describes the effect on [[IMAGE:173306d6b8cbd418_12_118]]
and [[IMAGE:173306d6b8cbd418_12_119]] ?















[[IMAGE:173306d6b8cbd418_12_120]] reduces to the conditional prediction [[IMAGE:173306d6b8cbd418_12_121]] alone, giving [[IMAGE:173306d6b8cbd418_12_122]] .



[[IMAGE:173306d6b8cbd418_12_123]] reduces to the unconditional prediction [[IMAGE:173306d6b8cbd418_12_124]] alone, giving [[IMAGE:173306d6b8cbd418_12_125]] .



[[IMAGE:173306d6b8cbd418_12_126]] reduces to the conditional prediction [[IMAGE:173306d6b8cbd418_12_127]] alone, giving [[IMAGE:173306d6b8cbd418_12_128]] .



[[IMAGE:173306d6b8cbd418_12_129]] remains unchanged from [[IMAGE:173306d6b8cbd418_12_130]] case because the guidance scale only
affects the output projection layer of the U-Net.


A published solution is not available for this question yet.
Question 19 NAT · 4.0 marks
Consider a 1-dimensional VAE with:
1. Encoder: [[IMAGE:173306d6b8cbd418_12_131]] , [[IMAGE:173306d6b8cbd418_12_132]] , [[IMAGE:173306d6b8cbd418_12_133]]
2. Decoder: [[IMAGE:173306d6b8cbd418_12_134]] , with fixed decoder variance [[IMAGE:173306d6b8cbd418_12_135]]
Use the reparameterization trick to express the latent sample [[IMAGE:173306d6b8cbd418_12_136]] . The
reconstruction loss [[IMAGE:173306d6b8cbd418_12_137]] for a single sample is given by [[IMAGE:173306d6b8cbd418_12_138]] .
Following are the values given: [[IMAGE:173306d6b8cbd418_12_139]] ,
input [[IMAGE:173306d6b8cbd418_12_140]] , sampled noise [[IMAGE:173306d6b8cbd418_12_141]] . Compute [[IMAGE:173306d6b8cbd418_12_142]] . Enter the answer correct to three
decimal places.












A published solution is not available for this question yet.
Question 20 MCQ · 1.0 marks
Inference in DDPM is much slower as compared to GAN or VAE.
True
False
A published solution is not available for this question yet.