da5007_2026T1_Q1_NA.pdf
Reinforcement Learning · Quiz 1 · Jan 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 2.0 marks
Which of the following is/are correct and valid reasons to consider sampling actions from a
softmax distribution instead of using an [[IMAGE:847d867b2fa2b5da_3_2]] greedy approach?
1.Under softmax exploration, the probability of selecting an action increases with its estimated
action value, which reduces unnecessary exploration of clearly inferior actions.
2.Unlike the [[IMAGE:847d867b2fa2b5da_3_3]] -greedy method, softmax exploration does not require careful, gradual decay of the
exploration parameter and still yields asymptotically correct behaviour even if the temperature is
reduced sharply.
3.It enables more fine-grained discrimination among actions whose estimated Q-values are close
to the maximum, allowing more nuanced preference for slightly better actions.
Which of the above statements is/are correct?


1, 2, 3
only 3
1, 2
1, 3
3, 2
Only 2
A published solution is not available for this question yet.
Question 3 MCQ · 2.0 marks
Consider a discounted return:
[[IMAGE:847d867b2fa2b5da_4_4]]
in an infinite-horizon MDP with bounded rewards [[IMAGE:847d867b2fa2b5da_4_5]] and discount factor [[IMAGE:847d867b2fa2b5da_4_6]]
1.For fixed rewards, the contribution of [[IMAGE:847d867b2fa2b5da_4_7]] to [[IMAGE:847d867b2fa2b5da_4_8]] decays geometrically as [[IMAGE:847d867b2fa2b5da_4_9]] .
2.If [[IMAGE:847d867b2fa2b5da_4_10]] and rewards are uniformly bounded, the infinite sum [[IMAGE:847d867b2fa2b5da_4_11]] is always finite.
3.For [[IMAGE:847d867b2fa2b5da_4_12]] , the discounted return [[IMAGE:847d867b2fa2b5da_4_13]] is bounded above in magnitude by [[IMAGE:847d867b2fa2b5da_4_14]] . Here, Rmax be
the maximum reward for any transition.
Which of the above statements is/are correct?











1, 2, 3
1, 3
2, 3
only 1
only 3
A published solution is not available for this question yet.
Question 4 MCQ · 1.0 marks
Consider the following assertion and reason pair and select the correct option:
**Assertion:**In the UCB algorithm for multi-armed bandits, replacing the upper confidence bound
with a lower confidence bound (and greedily selecting actions based on that lower bound) would
still promote effective exploration or lead to optimal reward maximisation.
**Reason:**Even when lower confidence bounds are used instead of upper bounds, if the algorithm
greedily selects arms based on these lower bounds, it would still encourage exploration and
ultimately achieve optimal reward maximisation.
Both Assertion and Reason are true, and Reason is the correct explanation of
Assertion.
Both Assertion and Reason are true, but Reason is NOT the correct
explanation of Assertion.
Assertion is true, Reason is false
Assertion is false, Reason is false
A published solution is not available for this question yet.
Question 5 MCQ · 1.0 marks
In the context of MDPs, what does it mean for a policy [[IMAGE:847d867b2fa2b5da_5_15]] to be greedy with respect to an action-
value function [[IMAGE:847d867b2fa2b5da_5_16]] ?


For each state [[IMAGE:847d867b2fa2b5da_5_17]] , [[IMAGE:847d867b2fa2b5da_5_18]] selects actions that minimise [[IMAGE:847d867b2fa2b5da_5_19]]



For each state [[IMAGE:847d867b2fa2b5da_5_20]] , [[IMAGE:847d867b2fa2b5da_5_21]] selects actions that maximise [[IMAGE:847d867b2fa2b5da_5_22]]



For each state [[IMAGE:847d867b2fa2b5da_5_23]] , [[IMAGE:847d867b2fa2b5da_5_24]] ignores [[IMAGE:847d867b2fa2b5da_5_25]] and follows a fixed, pre-defined schedule.



For each state [[IMAGE:847d867b2fa2b5da_5_26]] , [[IMAGE:847d867b2fa2b5da_5_27]] selects actions uniformly at random regardless of [[IMAGE:847d867b2fa2b5da_5_28]]



A published solution is not available for this question yet.
Question 6 MSQ · 2.0 marks
In policy iteration for finite MDPs, how are the Bellman equations used during the policy
evaluation step?
They are ignored; policy iteration does not rely on Bellman equations.
They are used to compute [[IMAGE:847d867b2fa2b5da_5_29]] exactly (or approximately) for the current policy
[[IMAGE:847d867b2fa2b5da_5_30]]


They are used to directly compute the optimal policy without evaluating
intermediate policies.
They are used to compute immediate rewards without considering transitions.
A published solution is not available for this question yet.
Question 7 MSQ · 2.0 marks
In a [[IMAGE:847d867b2fa2b5da_6_31]] -armed bandit setting, we maintain a running estimate of the action-value for each arm [[IMAGE:847d867b2fa2b5da_6_32]] ,
denoted [[IMAGE:847d867b2fa2b5da_6_33]] . We now consider using an upper confidence bound (UCB) style rule for arm
selection at time [[IMAGE:847d867b2fa2b5da_6_34]] , where [[IMAGE:847d867b2fa2b5da_6_35]] is the number of times arm [[IMAGE:847d867b2fa2b5da_6_36]] has been selected up to time [[IMAGE:847d867b2fa2b5da_6_37]] , and
[[IMAGE:847d867b2fa2b5da_6_38]] is a tunable hyperparameter.
Which of the following are good strategies for arm selection?








[[IMAGE:847d867b2fa2b5da_6_39]]

[[IMAGE:847d867b2fa2b5da_6_40]]

[[IMAGE:847d867b2fa2b5da_6_41]]

[[IMAGE:847d867b2fa2b5da_6_42]]

None of these
A published solution is not available for this question yet.
Question 8 MSQ · 2.0 marks
Which of the following equations best represents the Bellman optimality equation for the optimal
state-value function [[IMAGE:847d867b2fa2b5da_6_43]]

[[IMAGE:847d867b2fa2b5da_6_44]]

[[IMAGE:847d867b2fa2b5da_6_45]]

[[IMAGE:847d867b2fa2b5da_7_46]]

[[IMAGE:847d867b2fa2b5da_7_47]]

A published solution is not available for this question yet.
Question 9 NAT · 2.0 marks
The sequence of rewards for a continuing task with [[IMAGE:847d867b2fa2b5da_7_48]] is given below:
[[IMAGE:847d867b2fa2b5da_7_49]]
Find the return [[IMAGE:847d867b2fa2b5da_7_50]] . Your answer should have exactly two places after the decimal point.



A published solution is not available for this question yet.
Question 10 NAT · 2.0 marks
You have a 3-armed bandit where each arm, when active, yields a Bernoulli reward with success
probabilities [[IMAGE:847d867b2fa2b5da_7_51]] , [[IMAGE:847d867b2fa2b5da_7_52]] , and [[IMAGE:847d867b2fa2b5da_7_53]] . However, arm 1 is active in any given round with
probability [[IMAGE:847d867b2fa2b5da_7_54]] , and otherwise it yields a reward of 0 regardless of its Bernoulli outcome. Arms 2
and 3 are always active. What is the expected reward from pulling arm 1 in a single round?




A published solution is not available for this question yet.
Question 11 NAT · 1.0 marks
Consider a Markov Decision Process (MDP) with three states [[IMAGE:847d867b2fa2b5da_8_55]] , [[IMAGE:847d867b2fa2b5da_8_56]] , and [[IMAGE:847d867b2fa2b5da_8_57]] , where [[IMAGE:847d867b2fa2b5da_8_58]] is a terminal
state and [[IMAGE:847d867b2fa2b5da_8_59]] . Each non-terminal state has a single available action, and the transitions are
deterministic with associated rewards: the agent moves from [[IMAGE:847d867b2fa2b5da_8_60]] to [[IMAGE:847d867b2fa2b5da_8_61]] and receives a reward of 1,
and from [[IMAGE:847d867b2fa2b5da_8_62]] to [[IMAGE:847d867b2fa2b5da_8_63]] and receives a reward of 2.
Let the discount factor be [[IMAGE:847d867b2fa2b5da_8_64]] .
Based on the above data, answer the given subquestions.
Assume the initial value estimates are [[IMAGE:847d867b2fa2b5da_8_65]] , [[IMAGE:847d867b2fa2b5da_8_66]] , and [[IMAGE:847d867b2fa2b5da_8_67]] . After performing
one synchronous policy evaluation update (under the fixed policy that follows the chain),
determine [[IMAGE:847d867b2fa2b5da_8_68]] and [[IMAGE:847d867b2fa2b5da_8_69]] . Compute the ratio [[IMAGE:847d867b2fa2b5da_8_70]]
















A published solution is not available for this question yet.
Question 12 NAT · 1.0 marks
Consider a Markov Decision Process (MDP) with three states [[IMAGE:847d867b2fa2b5da_8_55]] , [[IMAGE:847d867b2fa2b5da_8_56]] , and [[IMAGE:847d867b2fa2b5da_8_57]] , where [[IMAGE:847d867b2fa2b5da_8_58]] is a terminal
state and [[IMAGE:847d867b2fa2b5da_8_59]] . Each non-terminal state has a single available action, and the transitions are
deterministic with associated rewards: the agent moves from [[IMAGE:847d867b2fa2b5da_8_60]] to [[IMAGE:847d867b2fa2b5da_8_61]] and receives a reward of 1,
and from [[IMAGE:847d867b2fa2b5da_8_62]] to [[IMAGE:847d867b2fa2b5da_8_63]] and receives a reward of 2.
Let the discount factor be [[IMAGE:847d867b2fa2b5da_8_64]] .
Based on the above data, answer the given subquestions.
Starting from [[IMAGE:847d867b2fa2b5da_9_71]] , [[IMAGE:847d867b2fa2b5da_9_72]] , and [[IMAGE:847d867b2fa2b5da_9_73]] . perform two successive synchronous
policy evaluation sweeps (under the fixed policy that follows the chain). What is [[IMAGE:847d867b2fa2b5da_9_74]] after these
two sweeps?














A published solution is not available for this question yet.
Question 13 NAT · 1.0 marks
Consider a 2-state MDP with states [[IMAGE:847d867b2fa2b5da_9_75]] and [[IMAGE:847d867b2fa2b5da_9_76]] , discount factor [[IMAGE:847d867b2fa2b5da_9_77]] . There are two actions in [[IMAGE:847d867b2fa2b5da_9_78]] :
[[IMAGE:847d867b2fa2b5da_9_79]] and [[IMAGE:847d867b2fa2b5da_9_80]] . Action [[IMAGE:847d867b2fa2b5da_9_81]] gives immediate reward [[IMAGE:847d867b2fa2b5da_9_82]] and moves deterministically to [[IMAGE:847d867b2fa2b5da_9_83]] ; action [[IMAGE:847d867b2fa2b5da_9_84]] gives
immediate reward [[IMAGE:847d867b2fa2b5da_9_85]] and keeps the agent in [[IMAGE:847d867b2fa2b5da_9_86]] . In [[IMAGE:847d867b2fa2b5da_9_87]] , there is a single action that yields an
immediate reward [[IMAGE:847d867b2fa2b5da_9_88]] and transitions back to [[IMAGE:847d867b2fa2b5da_9_89]] .
Assume initial value estimates are [[IMAGE:847d867b2fa2b5da_9_90]] , [[IMAGE:847d867b2fa2b5da_9_91]]
Based on the above data, answer the given subquestions.
If you perform a single value iteration Bellman optimality update for [[IMAGE:847d867b2fa2b5da_9_92]] , what is the new value
[[IMAGE:847d867b2fa2b5da_9_93]] ?



















A published solution is not available for this question yet.
Question 14 NAT · 1.0 marks
Consider a 2-state MDP with states [[IMAGE:847d867b2fa2b5da_9_75]] and [[IMAGE:847d867b2fa2b5da_9_76]] , discount factor [[IMAGE:847d867b2fa2b5da_9_77]] . There are two actions in [[IMAGE:847d867b2fa2b5da_9_78]] :
[[IMAGE:847d867b2fa2b5da_9_79]] and [[IMAGE:847d867b2fa2b5da_9_80]] . Action [[IMAGE:847d867b2fa2b5da_9_81]] gives immediate reward [[IMAGE:847d867b2fa2b5da_9_82]] and moves deterministically to [[IMAGE:847d867b2fa2b5da_9_83]] ; action [[IMAGE:847d867b2fa2b5da_9_84]] gives
immediate reward [[IMAGE:847d867b2fa2b5da_9_85]] and keeps the agent in [[IMAGE:847d867b2fa2b5da_9_86]] . In [[IMAGE:847d867b2fa2b5da_9_87]] , there is a single action that yields an
immediate reward [[IMAGE:847d867b2fa2b5da_9_88]] and transitions back to [[IMAGE:847d867b2fa2b5da_9_89]] .
Assume initial value estimates are [[IMAGE:847d867b2fa2b5da_9_90]] , [[IMAGE:847d867b2fa2b5da_9_91]]
Based on the above data, answer the given subquestions.
Suppose after some value iterations your current estimates are [[IMAGE:847d867b2fa2b5da_10_94]] , [[IMAGE:847d867b2fa2b5da_10_95]] . If you
perform one value iteration update for [[IMAGE:847d867b2fa2b5da_10_96]] , what is the new value [[IMAGE:847d867b2fa2b5da_10_97]] ?





















A published solution is not available for this question yet.
Question 15 MCQ · 2.0 marks
Consider an N-armed bandit problem in which each arm [[IMAGE:847d867b2fa2b5da_10_98]] , for [[IMAGE:847d867b2fa2b5da_10_99]] , produces a
reward of 1 with probability [[IMAGE:847d867b2fa2b5da_10_100]] and [[IMAGE:847d867b2fa2b5da_10_101]] otherwise, where [[IMAGE:847d867b2fa2b5da_10_102]] denotes the arm’s base success
probability.
Assume the bandit machine suffers from a hardware malfunction: when the agent attempts to pull
arm [[IMAGE:847d867b2fa2b5da_10_103]] , the intended action is not always executed. Instead, with probability [[IMAGE:847d867b2fa2b5da_10_104]] a wiring fault causes
the machine to randomly activate one arm chosen uniformly from the set [[IMAGE:847d867b2fa2b5da_10_105]] ,
regardless of the agent’s selection, and the reward is generated according to that arm’s reward
distribution. We refer to this malfunction as [[IMAGE:847d867b2fa2b5da_10_106]] -noise activation
Based on the above data, answer the given subquestions.
Consider a faulty [[IMAGE:847d867b2fa2b5da_10_107]] -armed bandit with [[IMAGE:847d867b2fa2b5da_10_108]] -noise activation and base arm success probabilities
[[IMAGE:847d867b2fa2b5da_11_109]] . If you always choose the arm with the highest base success probability [[IMAGE:847d867b2fa2b5da_11_110]] ,
what is the expected payoff per pull under this greedy strategy?













[[IMAGE:847d867b2fa2b5da_11_111]]

[[IMAGE:847d867b2fa2b5da_11_112]]

[[IMAGE:847d867b2fa2b5da_11_113]]

[[IMAGE:847d867b2fa2b5da_11_114]]

A published solution is not available for this question yet.
Question 16 MCQ · 2.0 marks
Consider an N-armed bandit problem in which each arm [[IMAGE:847d867b2fa2b5da_10_98]] , for [[IMAGE:847d867b2fa2b5da_10_99]] , produces a
reward of 1 with probability [[IMAGE:847d867b2fa2b5da_10_100]] and [[IMAGE:847d867b2fa2b5da_10_101]] otherwise, where [[IMAGE:847d867b2fa2b5da_10_102]] denotes the arm’s base success
probability.
Assume the bandit machine suffers from a hardware malfunction: when the agent attempts to pull
arm [[IMAGE:847d867b2fa2b5da_10_103]] , the intended action is not always executed. Instead, with probability [[IMAGE:847d867b2fa2b5da_10_104]] a wiring fault causes
the machine to randomly activate one arm chosen uniformly from the set [[IMAGE:847d867b2fa2b5da_10_105]] ,
regardless of the agent’s selection, and the reward is generated according to that arm’s reward
distribution. We refer to this malfunction as [[IMAGE:847d867b2fa2b5da_10_106]] -noise activation
Based on the above data, answer the given subquestions.
Consider a faulty [[IMAGE:847d867b2fa2b5da_11_115]] -armed bandit with [[IMAGE:847d867b2fa2b5da_11_116]] -noise activation and distinct base success probabilities
[[IMAGE:847d867b2fa2b5da_11_117]] . Suppose the agent follows a greedy strategy that always chooses the arm with the
highest empirical mean reward. How does the presence of activation noise affect the asymptotic
fraction of time for which the optimal arm (with success probability [[IMAGE:847d867b2fa2b5da_11_118]] ) is actually executed?













Noise has no effect; the optimal arm is activated almost surely in the long run.
Noise reduces the frequency with which the optimal arm is selected, but once
it is selected, it is always activated.
Noise eventually causes the greedy policy to oscillate indefinitely between
arms, preventing convergence.
Even when the greedy policy converges to always selecting the optimal arm,
activation noise prevents the optimal arm from being activated more than a fraction [[IMAGE:847d867b2fa2b5da_11_119]] of the
time.

A published solution is not available for this question yet.
Question 17 MCQ · 2.0 marks
Consider an N-armed bandit problem in which each arm [[IMAGE:847d867b2fa2b5da_10_98]] , for [[IMAGE:847d867b2fa2b5da_10_99]] , produces a
reward of 1 with probability [[IMAGE:847d867b2fa2b5da_10_100]] and [[IMAGE:847d867b2fa2b5da_10_101]] otherwise, where [[IMAGE:847d867b2fa2b5da_10_102]] denotes the arm’s base success
probability.
Assume the bandit machine suffers from a hardware malfunction: when the agent attempts to pull
arm [[IMAGE:847d867b2fa2b5da_10_103]] , the intended action is not always executed. Instead, with probability [[IMAGE:847d867b2fa2b5da_10_104]] a wiring fault causes
the machine to randomly activate one arm chosen uniformly from the set [[IMAGE:847d867b2fa2b5da_10_105]] ,
regardless of the agent’s selection, and the reward is generated according to that arm’s reward
distribution. We refer to this malfunction as [[IMAGE:847d867b2fa2b5da_10_106]] -noise activation
Based on the above data, answer the given subquestions.
In a faulty [[IMAGE:847d867b2fa2b5da_12_120]] -armed bandit with [[IMAGE:847d867b2fa2b5da_12_121]] -noise activation, suppose you mistakenly model the
environment as a standard bandit without activation noise. You estimate each arm’s success
probability using sample averages of observed rewards. How does the [[IMAGE:847d867b2fa2b5da_12_122]] -noise affect your
estimates of the true base success probabilities [[IMAGE:847d867b2fa2b5da_12_123]] ?













Your estimates are biased toward the overall average success probability
across all arms.
Your estimates are biased upward, overestimating all [[IMAGE:847d867b2fa2b5da_12_124]]

Your estimates remain unbiased for [[IMAGE:847d867b2fa2b5da_12_125]] because you still observe correct
rewards for the arm you believe you pulled.

Your estimates are biased downward, underestimating all [[IMAGE:847d867b2fa2b5da_12_126]]

A published solution is not available for this question yet.
Question 18 MCQ · 2.0 marks
Consider a large-scale news recommendation platform that uses a contextual bandit approach to
decide which headline to show to a user visiting the homepage. The context includes user features
(e.g., location, device type, time of day, coarse interest profile) and article features (e.g., topic,
recency, source reputation). Actions correspond to recommending one of the [[IMAGE:847d867b2fa2b5da_12_127]] candidate
headlines, and the reward is [[IMAGE:847d867b2fa2b5da_12_128]] if the user clicks the recommended headline and [[IMAGE:847d867b2fa2b5da_12_129]] otherwise.
Let the context vector [[IMAGE:847d867b2fa2b5da_12_130]] at time [[IMAGE:847d867b2fa2b5da_12_131]] encode both user and article features for the current
recommendation opportunity. There are [[IMAGE:847d867b2fa2b5da_12_132]] discrete actions, each corresponding to selecting one
of the candidate headlines for display.
Based on the above data, answer the given subquestions.
Which of the following best explains why a contextual bandit approach is preferred over a
standard multi-armed bandit in this news recommendation scenario?






Because user preferences are identical for all users, so context does not
influence click behaviour.
Because contextual bandits can tailor recommendations to each user and
article pair using the context, thereby improving click-through rates across heterogeneous users.
Because standard multi-armed bandits cannot perform any exploration at all.
Because contextual bandits guarantee that the same headline is always
optimal for every user.
A published solution is not available for this question yet.
Question 19 MCQ · 2.0 marks
Consider a large-scale news recommendation platform that uses a contextual bandit approach to
decide which headline to show to a user visiting the homepage. The context includes user features
(e.g., location, device type, time of day, coarse interest profile) and article features (e.g., topic,
recency, source reputation). Actions correspond to recommending one of the [[IMAGE:847d867b2fa2b5da_12_127]] candidate
headlines, and the reward is [[IMAGE:847d867b2fa2b5da_12_128]] if the user clicks the recommended headline and [[IMAGE:847d867b2fa2b5da_12_129]] otherwise.
Let the context vector [[IMAGE:847d867b2fa2b5da_12_130]] at time [[IMAGE:847d867b2fa2b5da_12_131]] encode both user and article features for the current
recommendation opportunity. There are [[IMAGE:847d867b2fa2b5da_12_132]] discrete actions, each corresponding to selecting one
of the candidate headlines for display.
Based on the above data, answer the given subquestions.
After many rounds of user interaction, what is the main optimality objective of the contextual
bandit algorithm in this recommendation setting?






Maximise the cumulative expected number of clicks by acting as closely as
possible to the best policy that maps contexts to headlines.
Ensure that on every single round, it always picks the empirically highest-CTR
headline observed so far, regardless of context.
Minimise the variance of rewards, even if that substantially reduces the total
number of clicks.
Focus only on the worst-performing user segment and ignore the rest.
A published solution is not available for this question yet.