da5002_2026T1_ET_FN.pdf
Mathematical Foundations of Generative AI · End Term · Jan 2026 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MSQ · 2.0 marks
Which of the following statements about the KL divergence is true?
When two distributions are identical, KL divergence between them is
maximized.
When two distributions are identical, KL divergence between them is
minimized.
KL divergence between any two distributions will always lie between 0 and 1.
Maximizing the ELBO for a Variational Autoencoder (VAE) involves minimizing
the KL divergence between the approximate posterior and the prior.
A published solution is not available for this question yet.
Question 3 MSQ · 2.0 marks
Consider a Generative Adversarial Network (GAN) trained to generate realistic images of cars.
Which of the following statements are true?
The generator tries to approximate the underlying distribution of car images.
The discriminator provides a probability that an input image is real or
generated.
At convergence, the discriminator is able to perfectly distinguish real and
generated car images.
The trained generator can produce novel car images that were not present in
the training dataset.
The generator directly maximizes the likelihood of the training data.
A published solution is not available for this question yet.
Question 4 NAT · 3.0 marks
Consider a GAN where the training dataset contains three times as many car images as bike
images. Suppose the generator has learned to produce only high-quality car images that are
indistinguishable from real ones. The discriminator is trained to output the probability that an
input image is real. What is the optimal output of the discriminator when given a car image? Enter
the answer correct to two decimal places.
A published solution is not available for this question yet.
Question 5 NAT · 3.0 marks
A Gaussian Mixture Model has two components, [[IMAGE:1c7e0624ff305b0d_3_2]] and [[IMAGE:1c7e0624ff305b0d_3_3]] , with prior probabilities
[[IMAGE:1c7e0624ff305b0d_3_4]] . The components are 1D Gaussians with parameters [[IMAGE:1c7e0624ff305b0d_3_5]] and
[[IMAGE:1c7e0624ff305b0d_3_6]] . For a data point [[IMAGE:1c7e0624ff305b0d_3_7]] , calculate the responsibility (posterior probability) of
component [[IMAGE:1c7e0624ff305b0d_3_8]] for this point, i.e., [[IMAGE:1c7e0624ff305b0d_3_9]] . Provide the answer to three decimal places.








A published solution is not available for this question yet.
Question 6 MCQ · 3.0 marks
Consider a VAE where we gradually increase the weight of the KL divergence term in the ELBO loss
function (i.e., increase the [[IMAGE:1c7e0624ff305b0d_4_10]] parameter in [[IMAGE:1c7e0624ff305b0d_4_11]] -VAE where the loss is
[[IMAGE:1c7e0624ff305b0d_4_12]] ). Assume a standard normal prior, [[IMAGE:1c7e0624ff305b0d_4_13]] . As
[[IMAGE:1c7e0624ff305b0d_4_14]] gradually increases from 0, what happens to the learned latent representations?





The distribution [[IMAGE:1c7e0624ff305b0d_4_15]] approaches [[IMAGE:1c7e0624ff305b0d_4_16]] but reconstructions become
worse.


The distribution [[IMAGE:1c7e0624ff305b0d_4_17]] goes farther from [[IMAGE:1c7e0624ff305b0d_4_18]] and reconstructions
become better.


The distribution [[IMAGE:1c7e0624ff305b0d_4_19]] approaches [[IMAGE:1c7e0624ff305b0d_4_20]] and reconstructions become
better.


The distribution [[IMAGE:1c7e0624ff305b0d_4_21]] goes farther from [[IMAGE:1c7e0624ff305b0d_4_22]] and reconstructions
become worse.


A published solution is not available for this question yet.
Question 7 MCQ · 2.0 marks
For a fixed input size and embedding dimension [[IMAGE:1c7e0624ff305b0d_4_23]] , consider standard multi-head attention where
the embedding dimension is evenly split across heads. Select the correct statements from the
following:

Query, key, and value projection layers in multi-head attention and single-
head attention contain an equal number of total learnable parameters.
Query, key, and value projection layers in multi-head attention contain more
learnable parameters than those in single-head attention.
Query, key, and value projection layers in multi-head attention contain fewer
learnable parameters than those in single-head attention.
A published solution is not available for this question yet.
Question 8 MSQ · 3.0 marks
In the context of the Transformer model's encoder-decoder architecture, which of the following
statements are accurate?
The encoder in a Transformer model is responsible for converting the input
sequence into a continuous representation that the decoder can then use to generate the output
sequence.
The decoder in a Transformer model only attends to the encoder's output
without any form of self-attention on its own inputs.
Multi-head attention in the encoder allows the model to jointly attend to
information from different representations at different positions.
The decoder's self-attention mechanism includes a masking component to
prevent attending to future positions, ensuring the model generates outputs one step at a time.
A published solution is not available for this question yet.
Question 9 NAT · 3.0 marks
A sequence of length [[IMAGE:1c7e0624ff305b0d_5_24]] is represented by
[[IMAGE:1c7e0624ff305b0d_5_25]]
where each row of [[IMAGE:1c7e0624ff305b0d_5_26]] represents a token. The attention module uses
[[IMAGE:1c7e0624ff305b0d_5_27]]
Thus
[[IMAGE:1c7e0624ff305b0d_5_28]]
Each row of [[IMAGE:1c7e0624ff305b0d_6_29]] and [[IMAGE:1c7e0624ff305b0d_6_30]] correspond to the representation of a single input token. Use [[IMAGE:1c7e0624ff305b0d_6_31]]
and scaling by [[IMAGE:1c7e0624ff305b0d_6_32]] .
Based on the above data, answer the given subquestions.
[[IMAGE:1c7e0624ff305b0d_6_34]]
Let the scaled scores for query [[IMAGE:1c7e0624ff305b0d_6_33]] be defined as . Compute the [[IMAGE:1c7e0624ff305b0d_6_35]] norm of
this vector. Enter the answer correct to two decimal places.












A published solution is not available for this question yet.
Question 10 MCQ · 3.0 marks
A sequence of length [[IMAGE:1c7e0624ff305b0d_5_24]] is represented by
[[IMAGE:1c7e0624ff305b0d_5_25]]
where each row of [[IMAGE:1c7e0624ff305b0d_5_26]] represents a token. The attention module uses
[[IMAGE:1c7e0624ff305b0d_5_27]]
Thus
[[IMAGE:1c7e0624ff305b0d_5_28]]
Each row of [[IMAGE:1c7e0624ff305b0d_6_29]] and [[IMAGE:1c7e0624ff305b0d_6_30]] correspond to the representation of a single input token. Use [[IMAGE:1c7e0624ff305b0d_6_31]]
and scaling by [[IMAGE:1c7e0624ff305b0d_6_32]] .
Based on the above data, answer the given subquestions.
Compute the attention output vector for the second token.









[[IMAGE:1c7e0624ff305b0d_6_36]]

[[IMAGE:1c7e0624ff305b0d_6_37]]

[[IMAGE:1c7e0624ff305b0d_6_38]]

[[IMAGE:1c7e0624ff305b0d_6_39]]

A published solution is not available for this question yet.
Question 11 MCQ · 1.0 marks
A decoder has sequence of length [[IMAGE:1c7e0624ff305b0d_7_40]] and pre-mask scaled score matrix
[[IMAGE:1c7e0624ff305b0d_7_41]]
Causal masking is applied to the scores exactly as in autoregressive transformers. The masking
uses
[[IMAGE:1c7e0624ff305b0d_7_42]]
for the mask matrix [[IMAGE:1c7e0624ff305b0d_7_43]] .Use
[[IMAGE:1c7e0624ff305b0d_7_44]]
Based on the above data, answer the given subquestions.
Which of the following is the masked score vector for the third query?





[[IMAGE:1c7e0624ff305b0d_7_45]]

[[IMAGE:1c7e0624ff305b0d_7_46]]

[[IMAGE:1c7e0624ff305b0d_7_47]]

[[IMAGE:1c7e0624ff305b0d_7_48]]

A published solution is not available for this question yet.
Question 12 NAT · 3.0 marks
A decoder has sequence of length [[IMAGE:1c7e0624ff305b0d_7_40]] and pre-mask scaled score matrix
[[IMAGE:1c7e0624ff305b0d_7_41]]
Causal masking is applied to the scores exactly as in autoregressive transformers. The masking
uses
[[IMAGE:1c7e0624ff305b0d_7_42]]
for the mask matrix [[IMAGE:1c7e0624ff305b0d_7_43]] .Use
[[IMAGE:1c7e0624ff305b0d_7_44]]
Based on the above data, answer the given subquestions.
For token 4, compute the attention probability assigned jointly to positions 1 and 2. Use the
masked attention weights. Enter the answer correct to three decimal places.





A published solution is not available for this question yet.
Question 13 MCQ · 4.0 marks
[[IMAGE:1c7e0624ff305b0d_8_49]]
Based on the above data, answer the given subquestions.
Find the mean and variance of [[IMAGE:1c7e0624ff305b0d_9_50]] .


Mean [[IMAGE:1c7e0624ff305b0d_9_51]] , Variance [[IMAGE:1c7e0624ff305b0d_9_52]]


Mean [[IMAGE:1c7e0624ff305b0d_9_53]] , Variance [[IMAGE:1c7e0624ff305b0d_9_54]]


Mean [[IMAGE:1c7e0624ff305b0d_9_55]] , Variance [[IMAGE:1c7e0624ff305b0d_9_56]]


Mean [[IMAGE:1c7e0624ff305b0d_9_57]] , Variance [[IMAGE:1c7e0624ff305b0d_9_58]]


A published solution is not available for this question yet.
Question 14 NAT · 3.0 marks
[[IMAGE:1c7e0624ff305b0d_8_49]]
Based on the above data, answer the given subquestions.
Compute the second coordinate of the final layer-normalized output. Enter the answer correct to
three decimal places.

A published solution is not available for this question yet.
Question 15 NAT · 2.0 marks
A language model is treated as a policy that generates a response of length 4. The rewards are
[[IMAGE:1c7e0624ff305b0d_9_59]]
with discount factor [[IMAGE:1c7e0624ff305b0d_9_60]] . Suppose the state-value estimates are
[[IMAGE:1c7e0624ff305b0d_10_61]]
The sampled log-policy gradient scalars [[IMAGE:1c7e0624ff305b0d_10_62]] are given as
[[IMAGE:1c7e0624ff305b0d_10_63]]
Based on the above data, answer the given subquestions.
Compute the full discounted return [[IMAGE:1c7e0624ff305b0d_10_64]] Enter the answer correct to
three decimal places.






A published solution is not available for this question yet.
Question 16 NAT · 2.0 marks
A language model is treated as a policy that generates a response of length 4. The rewards are
[[IMAGE:1c7e0624ff305b0d_9_59]]
with discount factor [[IMAGE:1c7e0624ff305b0d_9_60]] . Suppose the state-value estimates are
[[IMAGE:1c7e0624ff305b0d_10_61]]
The sampled log-policy gradient scalars [[IMAGE:1c7e0624ff305b0d_10_62]] are given as
[[IMAGE:1c7e0624ff305b0d_10_63]]
Based on the above data, answer the given subquestions.
Compute the reward-to-go [[IMAGE:1c7e0624ff305b0d_10_65]] . Enter the answer correct to two decimal places.






A published solution is not available for this question yet.
Question 17 NAT · 3.0 marks
A language model is treated as a policy that generates a response of length 4. The rewards are
[[IMAGE:1c7e0624ff305b0d_9_59]]
with discount factor [[IMAGE:1c7e0624ff305b0d_9_60]] . Suppose the state-value estimates are
[[IMAGE:1c7e0624ff305b0d_10_61]]
The sampled log-policy gradient scalars [[IMAGE:1c7e0624ff305b0d_10_62]] are given as
[[IMAGE:1c7e0624ff305b0d_10_63]]
Based on the above data, answer the given subquestions.
Compute the gradient estimate
[[IMAGE:1c7e0624ff305b0d_11_66]]
where [[IMAGE:1c7e0624ff305b0d_11_67]] is the advantage estimate given by [[IMAGE:1c7e0624ff305b0d_11_68]] Enter the answer correct to
three decimal places.








A published solution is not available for this question yet.
Question 18 NAT · 1.0 marks
[[IMAGE:1c7e0624ff305b0d_11_69]]
Based on the above data, answer the given subquestions.
Compute the importance sampling ratio [[IMAGE:1c7e0624ff305b0d_12_70]] for time step [[IMAGE:1c7e0624ff305b0d_12_71]] . Enter the answer correct to one
decimal place.



A published solution is not available for this question yet.
Question 19 NAT · 2.0 marks
[[IMAGE:1c7e0624ff305b0d_11_69]]
Based on the above data, answer the given subquestions.
Compute the clipped surrogate contribution at time 0. Enter the answer correct to one decimal
place.

A published solution is not available for this question yet.
Question 20 NAT · 2.0 marks
[[IMAGE:1c7e0624ff305b0d_11_69]]
Based on the above data, answer the given subquestions.
Compute the clipped surrogate contribution at time 1. Enter the answer correct to two decimal
places.

A published solution is not available for this question yet.
Question 21 NAT · 2.0 marks
[[IMAGE:1c7e0624ff305b0d_11_69]]
Based on the above data, answer the given subquestions.
Compute the clipped surrogate contribution at time 2. Enter the answer correct to one decimal
place.

A published solution is not available for this question yet.
Question 22 NAT · 1.0 marks
[[IMAGE:1c7e0624ff305b0d_11_69]]
Based on the above data, answer the given subquestions.
Compute the total PPO clipped objective contribution over the three time steps. Enter the answer
correct to two decimal places.

A published solution is not available for this question yet.