da5007_2025T3_Q1_NA.pdf
Reinforcement Learning · Quiz 1 · Sep 2025
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 235 SHORT_TEXT · 1.5 marks
Enter the correct answer for Blank (b) ________________
**NOTE:** Enter the exact answer without any extra space in the beginning or at the end.
A published solution is not available for this question yet.
Question 237 MCQ · 3.0 marks
[[IMAGE:8e32aab7e59b4517_2_0]]

[[IMAGE:8e32aab7e59b4517_2_1]]

[[IMAGE:8e32aab7e59b4517_2_2]]

[[IMAGE:8e32aab7e59b4517_2_3]]

[[IMAGE:8e32aab7e59b4517_2_4]]

A published solution is not available for this question yet.
Question 238 MCQ · 3.0 marks
Consider following assertion reason pair:
**Assertion:**
In the UCB algorithm for multi-armed bandits, replacing the upper confidence bound with a lower
confidence bound (and greedily selecting actions based on that lower bound)
would still promote effective exploration or lead to optimal reward maximisation.
**Reason:**
Even when lower confidence bounds are used instead of upper bounds, if the algorithm greedily
selects arms based on these lower bounds, it would still encourage exploration and ultimately
achieve optimal reward maximisation.
Both Assertion and Reason are true, and Reason is the correct explanation of
Assertion.
Both Assertion and Reason are true, but Reason is NOT the correct
explanation of Assertion.
Assertion is true, Reason is false
Assertion is false, Reason is false
A published solution is not available for this question yet.
Question 239 MCQ · 5.0 marks
[[IMAGE:8e32aab7e59b4517_3_5]]

Policy Iteration
Value Iteration
Both require equal updates
It cannot be determined
A published solution is not available for this question yet.
Question 240 MCQ · 2.0 marks
[[IMAGE:8e32aab7e59b4517_4_6]]

Arm 1
Arm 2
Arm 3
Arm 4
A published solution is not available for this question yet.
Question 241 MSQ · 5.0 marks
Which of the following statements correctly describes the differences in computational complexity
and convergence behaviour between Policy Iteration and Value Iteration algorithms in solving
Markov Decision Processes?
Value Iteration requires fewer iterations to converge, but each iteration is
computationally more expensive due to the policy evaluation step.
Value Iteration performs a combined update of value estimation and policy
improvement in one step and requires more iterations to converge compared to Policy Iteration.
Policy Iteration can only be applied to small state spaces, while Value Iteration
scales well to large state spaces.
Value Iteration is simpler to implement because it only maintains a value
function, while Policy Iteration maintains both policy and value function.
Both Policy Iteration and Value Iteration are guaranteed to converge to the
optimal policy.
A published solution is not available for this question yet.
Question 242 NAT · 5.0 marks
[[IMAGE:8e32aab7e59b4517_5_7]]

A published solution is not available for this question yet.
Question 243 NAT · 4.0 marks
[[IMAGE:8e32aab7e59b4517_6_8]]

A published solution is not available for this question yet.
Question 244 MCQ · 1.0 marks
Suppose you face a 2-armed bandit task where, at each time step, the true action values are either
(10, 20) with probability 0.7 (case A) or (90, 80) with probability 0.3 (case B).
Based on the above data, answer the given subquestions.
If you cannot observe which case you face, what is the best expected reward per step you can
achieve, and what strategy should you follow?
Randomly choose between Action 1 and Action 2; expected reward is 3.
Choose Action 1 always; expected reward is 34.
Choose Action 2 always; expected reward is 38.
Alternate between Action 1 and Action 2; expected reward is 40.
A published solution is not available for this question yet.
Question 245 MCQ · 1.0 marks
Suppose you face a 2-armed bandit task where, at each time step, the true action values are either
(10, 20) with probability 0.7 (case A) or (90, 80) with probability 0.3 (case B).
Based on the above data, answer the given subquestions.
If, at each time step, you are told whether you face case A or case B (but not the true action
values), what is the best expected reward per step you can achieve,and what strategy should you
follow?
Pick Action 2 in case A, Action 1 in case B; expected reward is 41.
Always pick Action 2; expected reward is 38.
Always pick Action 1; expected reward is 34.
The expected reward stated in all the given options is incorrect.
A published solution is not available for this question yet.
Question 246 MCQ · 2.0 marks
[[IMAGE:8e32aab7e59b4517_8_9]]
Based on the above data, answer the given subquestions.
Which of the following best explains why a contextual bandit approach is preferred over a
standard multi-armed bandit in this scenario?

Because the optimal waiting time is the same for all VMs, regardless of their
context.
Because contextual bandits can adaptively select actions based on the specific
features of each VM failure event, maximising expected reward across diverse situations.
Because multi-armed bandits are unable to balance exploration and
exploitation.
Because contextual bandits always guarantee zero regret after every round.
A published solution is not available for this question yet.
Question 247 MCQ · 4.0 marks
[[IMAGE:8e32aab7e59b4517_8_9]]
Based on the above data, answer the given subquestions.
[[IMAGE:8e32aab7e59b4517_9_10]]


It determines the penalty for rebooting or migrating the VM.
It is used to predict the expected reward for each possible waiting
time,allowing the algorithm to personalize decisions for each failure event.
It is ignored by LinUCB, which only uses past rewards.
It is used to randomly select an action to ensure exploration.
A published solution is not available for this question yet.
Question 248 MCQ · 2.0 marks
[[IMAGE:8e32aab7e59b4517_8_9]]
Based on the above data, answer the given subquestions.
What is the main optimality goal of a contextual bandit algorithm over many rounds?

Maximise the total expected reward by acting as closely as possible to the best
policy mapping contexts to actions.
Achieve zero regret after every individual round by always picking the
empirically best action so far for each context
Always exploit the highest-reward action seen so far
Minimise the loss on the worst round
A published solution is not available for this question yet.
Question 249 NAT · 2.0 marks
[[IMAGE:8e32aab7e59b4517_10_11]]
Based on the above data, answer the given subquestions.
What is the total reward if the robot collects both treasures and then exits,taking the shortest
possible path and never hitting a wall?

A published solution is not available for this question yet.
Question 250 MSQ · 3.0 marks
[[IMAGE:8e32aab7e59b4517_10_11]]
Based on the above data, answer the given subquestions.
[[IMAGE:8e32aab7e59b4517_11_12]]


[[IMAGE:8e32aab7e59b4517_11_13]]

[[IMAGE:8e32aab7e59b4517_11_14]]

[[IMAGE:8e32aab7e59b4517_11_15]]

[[IMAGE:8e32aab7e59b4517_11_16]]

A published solution is not available for this question yet.
Question 251 MCQ · 4.0 marks
Consider a pole-balancing task where the goal is to apply forces to a cart moving along a track to
keep a hinged pole from falling. A failure event occurs when the pole falls past a certain angle or
the cart moves off the track, after which the pole resets to vertical.
[[IMAGE:8e32aab7e59b4517_12_17]]
The problem can be formulated either as an episodic task (with episodes ending at failure) or as a
continuing task (after each failure, pole gets back into its initial position,
potentially infinite time horizon with discounting).Rewards can be defined as:
**• Episodic:** Reward +1 for each timestamp without failure, so the return is the number of steps
until failure (possibly infinite if balanced forever).
**• Continuing:** Reward 0 for each timestamp except at each failure, where the reward is −1.
Returns are discounted sums of future rewards.
Based on the above data, answer the given subquestions.
[[IMAGE:8e32aab7e59b4517_12_18]]


[[IMAGE:8e32aab7e59b4517_12_19]]

[[IMAGE:8e32aab7e59b4517_12_20]]

[[IMAGE:8e32aab7e59b4517_12_21]]

A published solution is not available for this question yet.