da5002_2026T2_Q1_NA.pdf
Mathematical Foundations of Generative AI · Quiz 1 · May 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MSQ · 1.0 marks
Let [[IMAGE:24d29b2d8e723982_2_2]] be a convex, left semi-continuous function with [[IMAGE:24d29b2d8e723982_2_3]] , and let [[IMAGE:24d29b2d8e723982_2_4]] , [[IMAGE:24d29b2d8e723982_2_5]] be two
probability distributions with densities defined on the same support [[IMAGE:24d29b2d8e723982_2_6]] . The [[IMAGE:24d29b2d8e723982_2_7]] -divergence is
defined as:
[[IMAGE:24d29b2d8e723982_2_8]]
Which of the following statements are correct?







[[IMAGE:24d29b2d8e723982_2_9]] is always symmetric, i.e. [[IMAGE:24d29b2d8e723982_2_10]] .


[[IMAGE:24d29b2d8e723982_2_11]] if and only if [[IMAGE:24d29b2d8e723982_2_12]] .


The KL-divergence [[IMAGE:24d29b2d8e723982_2_13]] is obtained by setting [[IMAGE:24d29b2d8e723982_2_14]] .


The KL-divergence [[IMAGE:24d29b2d8e723982_2_15]] is obtained by setting [[IMAGE:24d29b2d8e723982_2_16]] .


A published solution is not available for this question yet.
Question 3 MSQ · 2.0 marks
Each of the following [[IMAGE:24d29b2d8e723982_3_17]] -divergences is obtained by a specific choice of generator [[IMAGE:24d29b2d8e723982_3_18]] .
Which pairing of generator and divergence is/are correct?


[[IMAGE:24d29b2d8e723982_3_19]] gives the Total Variation distance.

[[IMAGE:24d29b2d8e723982_3_20]] gives the KL-divergence.

[[IMAGE:24d29b2d8e723982_3_21]] gives the (forward) KL-divergence.

[[IMAGE:24d29b2d8e723982_3_22]] gives the Jensen--Shannon divergence.

[[IMAGE:24d29b2d8e723982_3_23]] gives Total Variation distance.

A published solution is not available for this question yet.
Question 4 MSQ · 2.0 marks
Assume GAN training has converged to the global optimum, i.e., [[IMAGE:24d29b2d8e723982_3_24]] . Which of the following
statements is correct? Select all that apply.

The generator's loss is exactly zero.
The discriminator outputs [[IMAGE:24d29b2d8e723982_3_25]] everywhere, but the generator's loss is not zero.

At convergence the discriminator outputs [[IMAGE:24d29b2d8e723982_3_26]] on real samples and [[IMAGE:24d29b2d8e723982_3_27]] on
generated samples.


At optima, [[IMAGE:24d29b2d8e723982_3_28]] .

A published solution is not available for this question yet.
Question 5 MSQ · 4.0 marks
Which among the following functions can be used as [[IMAGE:24d29b2d8e723982_4_29]] -divergence generators? Select all that
apply.

[[IMAGE:24d29b2d8e723982_4_30]]

[[IMAGE:24d29b2d8e723982_4_31]]

[[IMAGE:24d29b2d8e723982_4_32]]

[[IMAGE:24d29b2d8e723982_4_33]]

A published solution is not available for this question yet.
Question 6 MSQ · 3.0 marks
Which of the following statements about the generator and discriminator in a Generative
Adversarial Network (GAN) are correct?
The discriminator is a binary classifier trained on a mix of real samples and
generated samples.
If the discriminator becomes too strong too quickly at telling real data apart
from the generator's fakes, the generator's gradients vanish and it stops learning.
The generator's goal is to maximize the discriminator's classification accuracy.
The discriminator minimizes the difference between real and generated data
distributions.
A published solution is not available for this question yet.
Question 7 MSQ · 3.0 marks
In a conditional GAN, both the generator [[IMAGE:24d29b2d8e723982_4_34]] and discriminator [[IMAGE:24d29b2d8e723982_4_35]] receive the
conditioning variable [[IMAGE:24d29b2d8e723982_4_36]] as additional input. Which of the following are correct statements about
cGANs?



At inference time, sampling [[IMAGE:24d29b2d8e723982_5_37]] for a fixed [[IMAGE:24d29b2d8e723982_5_38]] and varying
[[IMAGE:24d29b2d8e723982_5_39]] produces diverse samples all belonging to class [[IMAGE:24d29b2d8e723982_5_40]] .




The cGAN objective is identical to the unconditional GAN objective when [[IMAGE:24d29b2d8e723982_5_41]] is
taken to be a constant for all samples.

A cGAN can only be trained if [[IMAGE:24d29b2d8e723982_5_42]] is a discrete (categorical) label; continuous
conditioning variables are not possible within this framework.

If the discriminator ignores [[IMAGE:24d29b2d8e723982_5_43]] entirely, the generator may learn to ignore [[IMAGE:24d29b2d8e723982_5_44]]
too, producing samples that are realistic but not class-consistent.


A published solution is not available for this question yet.
Question 8 MCQ · 3.0 marks
For a convex function [[IMAGE:24d29b2d8e723982_5_45]] , the convex (Fenchel) conjugate is given by
[[IMAGE:24d29b2d8e723982_5_46]]
Which of the following is the Fenchel conjugate [[IMAGE:24d29b2d8e723982_5_47]] of the total variation distance?



[[IMAGE:24d29b2d8e723982_5_48]] for all [[IMAGE:24d29b2d8e723982_5_49]] .


[[IMAGE:24d29b2d8e723982_5_50]] for all [[IMAGE:24d29b2d8e723982_5_51]] .


[[IMAGE:24d29b2d8e723982_5_52]]

[[IMAGE:24d29b2d8e723982_5_53]] for all [[IMAGE:24d29b2d8e723982_5_54]] .


A published solution is not available for this question yet.
Question 9 MCQ · 3.0 marks
Consider the standard GAN objective from the classifier-guided perspective:
[[IMAGE:24d29b2d8e723982_6_55]]
Suppose the generator is fixed and the discriminator is trained to perfect accuracy. The
discriminator acheive the optimum at [[IMAGE:24d29b2d8e723982_6_56]] for the fixed generator. What is the gradient of [[IMAGE:24d29b2d8e723982_6_57]]
with respect to [[IMAGE:24d29b2d8e723982_6_58]] in the region where [[IMAGE:24d29b2d8e723982_6_59]] ?





A large positive gradient, pushing [[IMAGE:24d29b2d8e723982_6_60]] toward [[IMAGE:24d29b2d8e723982_6_61]] .


A well-defined gradient that decreases monotonically toward zero.
A zero or undefined gradient, providing no useful learning signal for the
generator.
[[IMAGE:24d29b2d8e723982_6_62]] .

A published solution is not available for this question yet.
Question 10 MCQ · 3.0 marks
In a CGAN trained on face dataset, the label [[IMAGE:24d29b2d8e723982_6_63]] is a length-2 one-hot vector: Blond [[IMAGE:24d29b2d8e723982_6_64]] vs. Not
Blond [[IMAGE:24d29b2d8e723982_6_65]] . As a controlled test, a single fixed noise vector [[IMAGE:24d29b2d8e723982_6_66]] is decoded twice - once with
[[IMAGE:24d29b2d8e723982_6_67]] and once with [[IMAGE:24d29b2d8e723982_6_68]] changing only the label. The two generated faces are nearly
identical in pose, identity, and background; only the hair colour differs. What does this best
demonstrate about a well-trained CGAN?






The label input [[IMAGE:24d29b2d8e723982_6_69]] has no real effect on the output; the result is purely due to
random chance.

The CGAN has decoupled the label attribute (hair colour) from the rest: [[IMAGE:24d29b2d8e723982_6_70]]
governs identity/pose/background, while [[IMAGE:24d29b2d8e723982_6_71]] independently governs the conditioned attribute.


The generator ignores [[IMAGE:24d29b2d8e723982_6_72]] once a label is provided, relying only on [[IMAGE:24d29b2d8e723982_6_73]] to generate
the entire image.


One-hot label vectors must always have length exactly 2 for a CGAN to work
correctly.
A published solution is not available for this question yet.
Question 11 MCQ · 3.0 marks
In a conditional GAN, the discriminator's outputs on four fake "car'' samples
(label [[IMAGE:24d29b2d8e723982_7_74]] ) are
[[IMAGE:24d29b2d8e723982_7_75]]
If we train the generator with the non-saturating loss [[IMAGE:24d29b2d8e723982_7_76]] , how do the
discriminator's outputs and the loss behave?



[[IMAGE:24d29b2d8e723982_7_77]] and [[IMAGE:24d29b2d8e723982_7_78]] increases.


[[IMAGE:24d29b2d8e723982_7_79]] and [[IMAGE:24d29b2d8e723982_7_80]] decreases toward [[IMAGE:24d29b2d8e723982_7_81]] , because the discriminator is
being fooled into rating fakes as real.



Both [[IMAGE:24d29b2d8e723982_7_82]] and [[IMAGE:24d29b2d8e723982_7_83]] stay constant.


[[IMAGE:24d29b2d8e723982_7_84]] and [[IMAGE:24d29b2d8e723982_7_85]] becomes negative.


A published solution is not available for this question yet.
Question 12 MCQ · 3.0 marks
A student proposes modifying a standard GAN so that, instead of feeding random noise [[IMAGE:24d29b2d8e723982_7_86]] into the
generator, they occasionally feed in real images and ask the generator to reproduce them. Which
of the following best describes the likely consequence of this modification?

It would have no effect, since the generator treats all inputs identically.
It shifts the generator's task toward reconstructing inputs, so it no longer
learns to sample novel images from the noise [[IMAGE:24d29b2d8e723982_7_87]] .

It guarantees the generator will perfectly memorize the training set.
It removes the need for a discriminator entirely.
A published solution is not available for this question yet.
Question 13 NAT · 3.0 marks
Let [[IMAGE:24d29b2d8e723982_8_88]] and [[IMAGE:24d29b2d8e723982_8_89]] be two probability distributions over [[IMAGE:24d29b2d8e723982_8_90]] :
[[IMAGE:24d29b2d8e723982_8_91]]
Compute [[IMAGE:24d29b2d8e723982_8_92]] . Use natural logarithm to give the answer. Enter your answer correct to two
decimal places.





A published solution is not available for this question yet.
Question 14 NAT · 3.0 marks
In a GAN, the real and generated distributions are mixtures of Dirac deltas on [[IMAGE:24d29b2d8e723982_8_93]] :
[[IMAGE:24d29b2d8e723982_8_94]]
Compute the GAN value function [[IMAGE:24d29b2d8e723982_8_95]] at the optimal discriminator [[IMAGE:24d29b2d8e723982_8_96]] . (Use natural
logarithm). Enter your answer correct to three decimal places.
Use: [[IMAGE:24d29b2d8e723982_8_97]] , [[IMAGE:24d29b2d8e723982_8_98]] , [[IMAGE:24d29b2d8e723982_8_99]] , [[IMAGE:24d29b2d8e723982_8_100]] .








A published solution is not available for this question yet.
Question 15 NAT · 3.0 marks
During the discriminator update step, the generator is held fixed and the discriminator is trained
to maximize
[[IMAGE:24d29b2d8e723982_9_101]]
with batch size [[IMAGE:24d29b2d8e723982_9_102]] . At one step, the discriminator outputs are:
[[IMAGE:24d29b2d8e723982_9_103]]
Use: [[IMAGE:24d29b2d8e723982_9_104]] , [[IMAGE:24d29b2d8e723982_9_105]] , [[IMAGE:24d29b2d8e723982_9_106]] , [[IMAGE:24d29b2d8e723982_9_107]] ,
[[IMAGE:24d29b2d8e723982_9_108]] .
Compute the value of [[IMAGE:24d29b2d8e723982_9_109]] for this mini-batch. Give your answer correct to three decimal places.









A published solution is not available for this question yet.
Question 16 NAT · 2.0 marks
Consider a CGAN for Fashion-MNIST (10 classes, grayscale images of size [[IMAGE:24d29b2d8e723982_9_110]] ). The critic
receives the [[IMAGE:24d29b2d8e723982_9_111]] image concatenated along the channel axis with 10 additional one-hot
label channels (each of spatial size [[IMAGE:24d29b2d8e723982_10_112]] ), giving a combined input tensor of shape
[[IMAGE:24d29b2d8e723982_10_113]] . The critic's first layer is a **Conv2D** with a [[IMAGE:24d29b2d8e723982_10_114]] kernel, no bias, mapping the
[[IMAGE:24d29b2d8e723982_10_115]] input to a [[IMAGE:24d29b2d8e723982_10_116]] output (stride 2). Compute the number of trainable
parameters in this layer.







A published solution is not available for this question yet.
Question 17 NAT · 2.0 marks
Consider two discrete distributions over [[IMAGE:24d29b2d8e723982_10_117]] :
[[IMAGE:24d29b2d8e723982_10_118]]
The ground distance is [[IMAGE:24d29b2d8e723982_10_119]] . Consider a transport plan [[IMAGE:24d29b2d8e723982_10_120]] , which is a [[IMAGE:24d29b2d8e723982_10_121]] matrix [[IMAGE:24d29b2d8e723982_10_122]]
with entries [[IMAGE:24d29b2d8e723982_10_123]] .
[[IMAGE:24d29b2d8e723982_10_124]]
Based on the above data, answer the given subquestions.
Compute the average work (total transport cost) [[IMAGE:24d29b2d8e723982_10_125]] of this plan. Enter the answer
correct to one decimal place.









A published solution is not available for this question yet.
Question 18 MCQ · 3.0 marks
Consider two discrete distributions over [[IMAGE:24d29b2d8e723982_10_117]] :
[[IMAGE:24d29b2d8e723982_10_118]]
The ground distance is [[IMAGE:24d29b2d8e723982_10_119]] . Consider a transport plan [[IMAGE:24d29b2d8e723982_10_120]] , which is a [[IMAGE:24d29b2d8e723982_10_121]] matrix [[IMAGE:24d29b2d8e723982_10_122]]
with entries [[IMAGE:24d29b2d8e723982_10_123]] .
[[IMAGE:24d29b2d8e723982_10_124]]
Based on the above data, answer the given subquestions.
A different valid plan between the same [[IMAGE:24d29b2d8e723982_11_126]] is given as
[[IMAGE:24d29b2d8e723982_11_127]]










[[IMAGE:24d29b2d8e723982_11_128]] is the optimal plan and [[IMAGE:24d29b2d8e723982_11_129]] is suboptimal.


[[IMAGE:24d29b2d8e723982_11_130]] achieves the minimum work ( [[IMAGE:24d29b2d8e723982_11_131]] ) and is therefore optimal, while [[IMAGE:24d29b2d8e723982_11_132]] is
valid but suboptimal.



Both plans are optimal.
Neither plan is valid.
A published solution is not available for this question yet.
Question 19 NAT · 3.0 marks
Inception features (assumed Gaussian) for real and generated image sets are 2-dimensional with
diagonal covariance matrices:
[[IMAGE:24d29b2d8e723982_11_133]]
Based on the above data, answer the given subquestions.
Compute the full FID score. Enter the answer correct to two decimal places.

A published solution is not available for this question yet.
Question 20 MCQ · 1.0 marks
Inception features (assumed Gaussian) for real and generated image sets are 2-dimensional with
diagonal covariance matrices:
[[IMAGE:24d29b2d8e723982_11_133]]
Based on the above data, answer the given subquestions.
What does a lower FID indicate?

Worse generative quality (generated features further from real features).
Better generative quality (generated features closer to real features).
A published solution is not available for this question yet.