da5007_2025T3_Q2_NA.pdf
Reinforcement Learning · Quiz 2 · Sep 2025
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 180 NAT · 0.5 marks
[[IMAGE:af9915b9b07d2db6_1_0]]

A published solution is not available for this question yet.
Question 182 MCQ · 2.0 marks
Consider the following assertion reason pair:
**Assertion**: In Monte Carlo (MC), the value function converges to the certainty equivalence
estimate.
**Reason**: Monte Carlo methods use complete trajectories to fit the value function as closely as
possible to the sampled returns.
Both Assertion and Reason are correct, and Reason is the correct explanation.
Assertion is correct, Reason is incorrect
Assertion is incorrect, Reason is correct
Both Assertion and Reason are correct, but Reason is not the correct
explanation.
A published solution is not available for this question yet.
Question 183 MSQ · 3.0 marks
Which of the following is the correct way to represent policy for policy search methods? Assume [[IMAGE:af9915b9b07d2db6_2_1]]
represents a real-valued parameter.

[[IMAGE:af9915b9b07d2db6_2_2]]

[[IMAGE:af9915b9b07d2db6_3_3]]

[[IMAGE:af9915b9b07d2db6_3_4]]

[[IMAGE:af9915b9b07d2db6_3_5]]

[[IMAGE:af9915b9b07d2db6_3_6]]

A published solution is not available for this question yet.
Question 184 MSQ · 3.0 marks
Recall the incremental update rule for REINFORCE:
[[IMAGE:af9915b9b07d2db6_3_7]]
Consider the following binary-bandit problem:
[[IMAGE:af9915b9b07d2db6_3_8]]
Which of the following expressions are equivalent to [[IMAGE:af9915b9b07d2db6_3_9]] ?



[[IMAGE:af9915b9b07d2db6_3_10]]

[[IMAGE:af9915b9b07d2db6_3_11]]

[[IMAGE:af9915b9b07d2db6_3_12]]

[[IMAGE:af9915b9b07d2db6_4_13]]

A published solution is not available for this question yet.
Question 185 MSQ · 2.0 marks
In Q-Learning, the update rule for the action-value function is based on bootstrapping from the
current estimate. Which of the following correctly represents this update?
[[IMAGE:af9915b9b07d2db6_4_14]]

[[IMAGE:af9915b9b07d2db6_4_15]]

[[IMAGE:af9915b9b07d2db6_4_16]]

[[IMAGE:af9915b9b07d2db6_4_17]]

A published solution is not available for this question yet.
Question 186 NAT · 3.0 marks
Consider an episodic task where the [[IMAGE:af9915b9b07d2db6_4_18]] -step returns follow a decaying exponential pattern:
[[IMAGE:af9915b9b07d2db6_4_19]]
Assume the episode is sufficiently long so the forward-view infinite sum is valid.
Compute the [[IMAGE:af9915b9b07d2db6_4_20]] -return for [[IMAGE:af9915b9b07d2db6_4_21]] . Enter your answer correct to **two decimal places**.




A published solution is not available for this question yet.
Question 187 MCQ · 2.0 marks
In Deep Q-Networks (DQN), a separate **target network** is maintained, whose weights are
periodically copied from the main (online) network instead of being updated at every step.
Based on the above data, answer the given subquestions.
What is the primary reason for using a separate target network with periodically updated weights
in DQN?
To accelerate convergence by increasing the learning rate.
To stabilize learning by keeping target values fixed for several steps.
To ensure the Q-values are always up to date.
To share weights between actor and critic models.
A published solution is not available for this question yet.
Question 188 MSQ · 2.0 marks
In Deep Q-Networks (DQN), a separate **target network** is maintained, whose weights are
periodically copied from the main (online) network instead of being updated at every step.
Based on the above data, answer the given subquestions.
What are the effects of periodically updating the weights of the target network in DQN?
Provides stable targets to the main network during training.
Causes the learning targets to be non-stationary.
Mitigates instabilities in Q-learning updates.
Avoids the need for a replay buffer.
A published solution is not available for this question yet.
Question 189 MCQ · 1.0 marks
[[IMAGE:af9915b9b07d2db6_6_22]]
Based on the above data, answer the given subquestions.
Find the Q values for all three actions [[IMAGE:af9915b9b07d2db6_6_23]] for the state [[IMAGE:af9915b9b07d2db6_6_24]] . Express your answer as a vector q,
where [[IMAGE:af9915b9b07d2db6_6_25]]




[[IMAGE:af9915b9b07d2db6_6_26]]

[[IMAGE:af9915b9b07d2db6_6_27]]

[[IMAGE:af9915b9b07d2db6_6_28]]

[[IMAGE:af9915b9b07d2db6_6_29]]

A published solution is not available for this question yet.
Question 190 NAT · 4.0 marks
[[IMAGE:af9915b9b07d2db6_6_22]]
Based on the above data, answer the given subquestions.
Compute the TD target for this transition using SARSA with [[IMAGE:af9915b9b07d2db6_7_30]] . Use a greedy policy derived
from the Q-values at [[IMAGE:af9915b9b07d2db6_7_31]] to select the next action.



A published solution is not available for this question yet.
Question 191 MCQ · 4.0 marks
[[IMAGE:af9915b9b07d2db6_6_22]]
Based on the above data, answer the given subquestions.
Perform one Step of semi-gradient TD using this transition. What would be [[IMAGE:af9915b9b07d2db6_7_32]] if [[IMAGE:af9915b9b07d2db6_7_33]] ?



[[IMAGE:af9915b9b07d2db6_7_34]]

[[IMAGE:af9915b9b07d2db6_7_35]]

[[IMAGE:af9915b9b07d2db6_7_36]]

[[IMAGE:af9915b9b07d2db6_7_37]]

A published solution is not available for this question yet.
Question 192 MCQ · 4.0 marks
Consider the following algorithm for learning action-value estimates in an episodic MDP:
[[IMAGE:af9915b9b07d2db6_8_38]]
Based on the above data, answer the given subquestions.
Which of the following issues in standard Q-learning does the given algorithm aim to reduce?

Overestimation of action values due to maximisation bias.
Underestimation of terminal state values.
Instability due to non-stationary rewards.
Bias introduced by [[IMAGE:af9915b9b07d2db6_8_39]] -greedy exploration.

A published solution is not available for this question yet.
Question 193 MSQ · 4.0 marks
Consider the following algorithm for learning action-value estimates in an episodic MDP:
[[IMAGE:af9915b9b07d2db6_8_38]]
Based on the above data, answer the given subquestions.
In the given pseudocode, the algorithm maintains two value functions, [[IMAGE:af9915b9b07d2db6_8_40]] and [[IMAGE:af9915b9b07d2db6_8_41]] , and updates
them alternately.
Which of the following best describes the purpose and effect of this design choice?



It ensures that the same target network is used for both selection and
evaluation, improving stability.
It ensures the overall estimate equals the true expected return by averaging
[[IMAGE:af9915b9b07d2db6_9_42]] and [[IMAGE:af9915b9b07d2db6_9_43]] .


It decorrelates the selection and evaluation of the greedy action, reducing
overestimation bias.
It doubles the effective learning rate by updating two estimators in parallel.
A published solution is not available for this question yet.
Question 194 MCQ · 3.0 marks
Recall from the lectures, the forward-view [[IMAGE:af9915b9b07d2db6_9_44]] -**return** is defined as:
[[IMAGE:af9915b9b07d2db6_9_45]]
where [[IMAGE:af9915b9b07d2db6_9_46]] denotes the [[IMAGE:af9915b9b07d2db6_9_47]] -step return starting at time [[IMAGE:af9915b9b07d2db6_9_48]] .Consider using the TD( [[IMAGE:af9915b9b07d2db6_9_49]] ) algorithm for an
**episodic** task.
Based on the above data, answer the given subquestions.
What is the effect of setting [[IMAGE:af9915b9b07d2db6_9_50]] in TD( [[IMAGE:af9915b9b07d2db6_9_51]] ) on the weighting of [[IMAGE:af9915b9b07d2db6_9_52]] -step returns?









All [[IMAGE:af9915b9b07d2db6_9_53]] -step returns are ignored.

Only the 1-step return contributes, with weight 1.
All [[IMAGE:af9915b9b07d2db6_9_54]] -step returns contribute equally.

The weighting distribution remains unchanged.
A published solution is not available for this question yet.
Question 195 MCQ · 3.0 marks
Recall from the lectures, the forward-view [[IMAGE:af9915b9b07d2db6_9_44]] -**return** is defined as:
[[IMAGE:af9915b9b07d2db6_9_45]]
where [[IMAGE:af9915b9b07d2db6_9_46]] denotes the [[IMAGE:af9915b9b07d2db6_9_47]] -step return starting at time [[IMAGE:af9915b9b07d2db6_9_48]] .Consider using the TD( [[IMAGE:af9915b9b07d2db6_9_49]] ) algorithm for an
**episodic** task.
Based on the above data, answer the given subquestions.
As [[IMAGE:af9915b9b07d2db6_10_55]] in an episodic TD( [[IMAGE:af9915b9b07d2db6_10_56]] ) task, the [[IMAGE:af9915b9b07d2db6_10_57]] -return [[IMAGE:af9915b9b07d2db6_10_58]] approaches which of the following?










The Monte Carlo return.
The one-step TD target.
The average of all [[IMAGE:af9915b9b07d2db6_10_59]] -step returns.

Zero, since weights vanish as [[IMAGE:af9915b9b07d2db6_10_60]] .

A published solution is not available for this question yet.
Question 196 NAT · 3.0 marks
Consider an episodic task with four non-terminal states: A, B, C and D. The following are some
episodes experienced by an agent following a fixed policy. The terminal state is not explicitly
mentioned for any of the episodes.
A,1, C,1, D,0, C,0, A,1, B,1
A,0, C,0, D,1, C,1, B,1
D,1, C,0, A,0, C,1, B,0
C,0, A,1, C,1, B,0
C,1, B,0
D,0, B,1
B,1
B,0
(Enter your answer correct to **three decimal places**.)
Based on the above data, answer the given subquestions.
Using every-visit Monte Carlo, estimate [[IMAGE:af9915b9b07d2db6_10_61]] and [[IMAGE:af9915b9b07d2db6_10_62]] . What is the value of [[IMAGE:af9915b9b07d2db6_10_63]] ?
(use [[IMAGE:af9915b9b07d2db6_10_64]] )




A published solution is not available for this question yet.