da5002_2025T3_ET_FN.pdf
Mathematical Foundations of Generative AI · End Term · Sep 2025 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 3.0 marks
In the context of Denoising Diffusion Probabilistic Models (DDPM), the trained noise predictor
[[IMAGE:ef8a71bf9c941dbd_2_2]] approximates the score function of the data distribution via
[[IMAGE:ef8a71bf9c941dbd_2_3]]
Consider a specific timestep [[IMAGE:ef8a71bf9c941dbd_2_4]] where the cumulative noise variance schedule is given by [[IMAGE:ef8a71bf9c941dbd_2_5]]
. At this timestep, for a specific input [[IMAGE:ef8a71bf9c941dbd_2_6]] , the model predicts a noise vector:
[[IMAGE:ef8a71bf9c941dbd_2_7]]
Calculate the estimated score vector, [[IMAGE:ef8a71bf9c941dbd_2_8]] .







[[IMAGE:ef8a71bf9c941dbd_2_9]]

[[IMAGE:ef8a71bf9c941dbd_2_10]]

[[IMAGE:ef8a71bf9c941dbd_2_11]]

[[IMAGE:ef8a71bf9c941dbd_2_12]]

A published solution is not available for this question yet.
Question 3 MCQ · 3.0 marks
You are performing inference using a Guided Diffusion model with Classifier-Free Guidance (CFG)
with a guidance scale [[IMAGE:ef8a71bf9c941dbd_3_13]] . At a specific step, the model outputs:
[[IMAGE:ef8a71bf9c941dbd_3_14]]
• Unconditional prediction:
[[IMAGE:ef8a71bf9c941dbd_3_16]]
• Conditional prediction (given class [[IMAGE:ef8a71bf9c941dbd_3_15]] ):
Using the standard CFG formula, calculate the final guided noise vector [[IMAGE:ef8a71bf9c941dbd_3_17]] .





[[IMAGE:ef8a71bf9c941dbd_3_18]]

[[IMAGE:ef8a71bf9c941dbd_3_19]]

[[IMAGE:ef8a71bf9c941dbd_3_20]]

[[IMAGE:ef8a71bf9c941dbd_3_21]]

A published solution is not available for this question yet.
Question 4 MCQ · 3.0 marks
In a Deep Transformer with Residual connections, the output of layer [[IMAGE:ef8a71bf9c941dbd_4_22]] is [[IMAGE:ef8a71bf9c941dbd_4_23]] .
Assume that at initialization, the function [[IMAGE:ef8a71bf9c941dbd_4_24]] (Attention + FFN) outputs values with variance
roughly [[IMAGE:ef8a71bf9c941dbd_4_25]] .
After 10 layers, assuming independence, approximately how much has the signal variance
increased relative to the input variance [[IMAGE:ef8a71bf9c941dbd_4_26]] ?
(Approximation: [[IMAGE:ef8a71bf9c941dbd_4_27]] ).






It remains the same [[IMAGE:ef8a71bf9c941dbd_4_28]] .

It increases by a factor of [[IMAGE:ef8a71bf9c941dbd_4_29]] .

It doubles [[IMAGE:ef8a71bf9c941dbd_4_30]] .

It increases 10-fold.
A published solution is not available for this question yet.
Question 5 MCQ · 3.0 marks
You are performing Beam Search with width [[IMAGE:ef8a71bf9c941dbd_4_31]] . At step [[IMAGE:ef8a71bf9c941dbd_4_32]] , you have two active beams with
accumulated log-probs: Beam A: [[IMAGE:ef8a71bf9c941dbd_4_33]] (ending in token 'cat') Beam B: [[IMAGE:ef8a71bf9c941dbd_4_34]] (ending in token 'dog')
At step [[IMAGE:ef8a71bf9c941dbd_4_35]] , the model gives the following top log-probs: Given 'cat': 'sits' (-0.5), 'naps' (-1.0)
Given 'dog': 'barks' (-0.1), 'runs' (-0.4) Which two paths form the new beams?





('cat', 'sits') and ('dog', 'barks')
('cat', 'sits') and ('cat', 'naps')
('dog', 'barks') and ('dog', 'runs')
('cat', 'naps') and ('dog', 'runs')
A published solution is not available for this question yet.
Question 6 MCQ · 3.0 marks
[[IMAGE:ef8a71bf9c941dbd_5_36]]

[[IMAGE:ef8a71bf9c941dbd_5_37]]

[[IMAGE:ef8a71bf9c941dbd_5_38]]

[[IMAGE:ef8a71bf9c941dbd_5_39]]

Undefined
A published solution is not available for this question yet.
Question 7 MCQ · 3.0 marks
When treating an Auto-Regressive Language Model as an RL policy:
• The **State** is the context (prompt + generated tokens so far).
• The **Action Space** is the vocabulary [[IMAGE:ef8a71bf9c941dbd_5_40]] .
If the vocabulary size is [[IMAGE:ef8a71bf9c941dbd_5_41]] and we generate a sequence of length [[IMAGE:ef8a71bf9c941dbd_5_42]] .
What is the dimension of the policy output distribution at a single timestep, and what is the
dimensionality of the trajectory space?



Output dim: [[IMAGE:ef8a71bf9c941dbd_5_43]] ; Trajectory space: [[IMAGE:ef8a71bf9c941dbd_5_44]] .


Output dim: [[IMAGE:ef8a71bf9c941dbd_5_45]] ; Trajectory space: [[IMAGE:ef8a71bf9c941dbd_5_46]] .


Output dim: 1; Trajectory space: [[IMAGE:ef8a71bf9c941dbd_5_47]] .

Output dim: 50,000; Trajectory space: [[IMAGE:ef8a71bf9c941dbd_5_48]] .

A published solution is not available for this question yet.
Question 8 MCQ · 3.0 marks
Trust Region Policy Optimization (TRPO) maximizes the surrogate objective subject to a constraint
on the KL divergence between the old and new policies. What is the nature of this constraint?
[[IMAGE:ef8a71bf9c941dbd_6_49]]

[[IMAGE:ef8a71bf9c941dbd_6_50]]

[[IMAGE:ef8a71bf9c941dbd_6_51]]

[[IMAGE:ef8a71bf9c941dbd_6_52]]

A published solution is not available for this question yet.
Question 9 MCQ · 3.0 marks
The Policy Gradient theorem states [[IMAGE:ef8a71bf9c941dbd_6_53]] . If we introduce a baseline
[[IMAGE:ef8a71bf9c941dbd_6_54]] that depends only on state, the term becomes [[IMAGE:ef8a71bf9c941dbd_6_55]] . What is the
expected value of the following baseline term? [[IMAGE:ef8a71bf9c941dbd_6_56]]




[[IMAGE:ef8a71bf9c941dbd_6_57]]

[[IMAGE:ef8a71bf9c941dbd_6_58]]

[[IMAGE:ef8a71bf9c941dbd_6_59]]

[[IMAGE:ef8a71bf9c941dbd_6_60]]

A published solution is not available for this question yet.
Question 10 MCQ · 3.0 marks
DDIMs are faster than DDPM mainly because:
They remove the neural network entirely.
They require more noise steps.
They require additional forward passes.
The sampling process is deterministic when [[IMAGE:ef8a71bf9c941dbd_7_61]] .

A published solution is not available for this question yet.
Question 11 MSQ · 4.0 marks
An auto-regressive language model predicts the next token [[IMAGE:ef8a71bf9c941dbd_7_62]] given the history [[IMAGE:ef8a71bf9c941dbd_7_63]] . In the RL
interpretation used for alignment, which of the following correspondences between LM objects
and RL objects are correct?


The state [[IMAGE:ef8a71bf9c941dbd_7_64]] corresponds to the current context: prompt plus all tokens
generated so far.

The action [[IMAGE:ef8a71bf9c941dbd_7_65]] corresponds to selecting the next token from the vocabulary.

The trajectory [[IMAGE:ef8a71bf9c941dbd_7_66]] corresponds to the full sequence of (prompt, generated
tokens) over time.

The reward [[IMAGE:ef8a71bf9c941dbd_7_67]] is necessarily the log-likelihood of the token under the
pretraining objective.

A published solution is not available for this question yet.
Question 12 MSQ · 4.0 marks
The sinusoidal positional encoding (PE) for Transformers is defined as:
[[IMAGE:ef8a71bf9c941dbd_8_68]]
Based on the properties of sinusoidal positional encodings, identify all correct statements:

Positional encodings allow the model to incorporate order information that
attention alone cannot provide.
The sinusoidal frequencies grow geometrically with the dimension index [[IMAGE:ef8a71bf9c941dbd_8_69]] .

Sinusoidal positional encodings are learned parameters updated during
training.
Adding positional encodings to token embeddings guarantees each position
has a unique representation.
A published solution is not available for this question yet.
Question 13 NAT · 3.0 marks
Consider a single-head self-attention mechanism with [[IMAGE:ef8a71bf9c941dbd_8_70]] . You are given the following Query (
[[IMAGE:ef8a71bf9c941dbd_8_71]] ) and Key ( [[IMAGE:ef8a71bf9c941dbd_8_72]] ) matrices for a sequence length of [[IMAGE:ef8a71bf9c941dbd_8_73]] :
[[IMAGE:ef8a71bf9c941dbd_8_74]]
Compute the raw attention scores (before Softmax) and scale them by the factor [[IMAGE:ef8a71bf9c941dbd_8_75]] . What is
the value of the element at position [[IMAGE:ef8a71bf9c941dbd_8_76]] (i.e., row 1, column 2) of the scaled score matrix?







A published solution is not available for this question yet.
Question 14 NAT · 3.0 marks
You have a Transformer with model dimension [[IMAGE:ef8a71bf9c941dbd_9_77]] . You employ [[IMAGE:ef8a71bf9c941dbd_9_78]] parallel
attention heads. The output of the heads is concatenated and projected by a matrix [[IMAGE:ef8a71bf9c941dbd_9_79]] . If the
input batch size is [[IMAGE:ef8a71bf9c941dbd_9_80]] and sequence length is [[IMAGE:ef8a71bf9c941dbd_9_81]] , what is the total number of floating
point elements in the final output tensor of the Multi-Head Attention block (before the residual
connection)?





A published solution is not available for this question yet.
Question 15 NAT · 3.0 marks
In RLHF (Reinforcement Learning from Human Feedback), the reward used to update the
language model [[IMAGE:ef8a71bf9c941dbd_9_82]] is defined as
[[IMAGE:ef8a71bf9c941dbd_9_83]]
Given:
• Reward model score [[IMAGE:ef8a71bf9c941dbd_9_84]]
• [[IMAGE:ef8a71bf9c941dbd_9_85]]
• [[IMAGE:ef8a71bf9c941dbd_9_86]]
• KL coefficient [[IMAGE:ef8a71bf9c941dbd_9_87]]
Calculate the total reward. (Use [[IMAGE:ef8a71bf9c941dbd_9_88]] .)







A published solution is not available for this question yet.
Question 16 NAT · 3.0 marks
You are training a Reward Model (RM) using the Bradley-Terry model.
For a specific prompt [[IMAGE:ef8a71bf9c941dbd_10_89]] , you have a winning completion [[IMAGE:ef8a71bf9c941dbd_10_90]] and a losing completion [[IMAGE:ef8a71bf9c941dbd_10_91]] . The RM
currently predicts the scalar rewards as:
[[IMAGE:ef8a71bf9c941dbd_10_92]]
Calculate the negative log-likelihood loss for this pair: [[IMAGE:ef8a71bf9c941dbd_10_93]] .





A published solution is not available for this question yet.
Question 17 NAT · 3.0 marks
A continuous time State-Space Model (SSM) has a scalar state parameter [[IMAGE:ef8a71bf9c941dbd_10_94]] (representing
decay). You discretize this system with a step size [[IMAGE:ef8a71bf9c941dbd_10_95]] using the Zero-Order Hold (ZOH)
method (or approximation thereom).
Calculate the discretized state transition parameter [[IMAGE:ef8a71bf9c941dbd_10_96]] .
Formula: [[IMAGE:ef8a71bf9c941dbd_10_97]] .




A published solution is not available for this question yet.