da5007_2023T2_Q1_NA.pdf
Reinforcement Learning · Quiz 1 · May 2023
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 137 MCQ · 3.0 marks
[[IMAGE:2128091633a81988_2_0]]

[[IMAGE:2128091633a81988_2_1]]

[[IMAGE:2128091633a81988_2_2]]

[[IMAGE:2128091633a81988_2_3]]

[[IMAGE:2128091633a81988_3_4]]

A published solution is not available for this question yet.
Question 138 MCQ · 3.0 marks
[[IMAGE:2128091633a81988_3_5]]

[[IMAGE:2128091633a81988_3_6]]

[[IMAGE:2128091633a81988_3_7]]

[[IMAGE:2128091633a81988_3_8]]

[[IMAGE:2128091633a81988_3_9]]

A published solution is not available for this question yet.
Question 139 MCQ · 3.0 marks
In the context of a multi-armed bandit problem with stationary reward distributions, consider the
following:
**Assertion:** UCB minimizes the regret better than the softmax approach.
**Reason:** Softmax approach assigns a low probability of picking a sub-optimal arm that has a very
low expected reward.
Assertion and Reason are both true and Reason is a correct explanation of the
Assertion.
Assertion and Reason are both true and Reason is not a correct explanation of
the Assertion.
Assertion is true and Reason is false
Both Assertion and Reason are false.
A published solution is not available for this question yet.
Question 140 MCQ · 2.0 marks
[[IMAGE:2128091633a81988_4_10]]

TRUE
FALSE
A published solution is not available for this question yet.
Question 141 MCQ · 4.0 marks
[[IMAGE:2128091633a81988_5_11]]

[[IMAGE:2128091633a81988_5_12]]

[[IMAGE:2128091633a81988_5_13]]

[[IMAGE:2128091633a81988_5_14]]

[[IMAGE:2128091633a81988_5_15]]

A published solution is not available for this question yet.
Question 142 MSQ · 4.0 marks
[[IMAGE:2128091633a81988_6_16]]

1
2
3
4
5
A published solution is not available for this question yet.
Question 143 MSQ · 4.0 marks
[[IMAGE:2128091633a81988_6_17]]

[[IMAGE:2128091633a81988_6_18]]

[[IMAGE:2128091633a81988_6_19]]

[[IMAGE:2128091633a81988_7_20]]

[[IMAGE:2128091633a81988_7_21]]

A published solution is not available for this question yet.
Question 144 MSQ · 3.0 marks
[[IMAGE:2128091633a81988_7_22]]

[[IMAGE:2128091633a81988_7_23]]

[[IMAGE:2128091633a81988_7_24]]

[[IMAGE:2128091633a81988_7_25]]

[[IMAGE:2128091633a81988_7_26]]

[[IMAGE:2128091633a81988_8_27]]

A published solution is not available for this question yet.
Question 145 NAT · 4.0 marks
Consider the following statements, all of which are regarding policy iteration run on a finite MDP:
(1) Policy iteration can be used to find a deterministic optimal policy.
(2) Policy iteration is an algorithm that is exclusively used to evaluate the value function for a given
policy.
(3) We can use the optimal value function output by policy iteration to find out all possible optimal
policies, both deterministic and stochastic.
How many of these statements are true?
A published solution is not available for this question yet.
Question 146 NAT · 4.0 marks
[[IMAGE:2128091633a81988_9_28]]

A published solution is not available for this question yet.
Question 147 NAT · 4.0 marks
[[IMAGE:2128091633a81988_9_29]]

A published solution is not available for this question yet.
Question 148 NAT · 1.5 marks
[[IMAGE:2128091633a81988_10_30]]
Based on the above data, answer the given subquestions.
[[IMAGE:2128091633a81988_10_31]]


A published solution is not available for this question yet.
Question 149 NAT · 1.5 marks
[[IMAGE:2128091633a81988_10_30]]
Based on the above data, answer the given subquestions.
[[IMAGE:2128091633a81988_11_32]]


A published solution is not available for this question yet.
Question 150 NAT · 1.5 marks
[[IMAGE:2128091633a81988_10_30]]
Based on the above data, answer the given subquestions.
[[IMAGE:2128091633a81988_11_33]]


A published solution is not available for this question yet.
Question 151 NAT · 1.5 marks
[[IMAGE:2128091633a81988_10_30]]
Based on the above data, answer the given subquestions.
[[IMAGE:2128091633a81988_12_34]]


A published solution is not available for this question yet.
Question 152 NAT · 1.0 marks
[[IMAGE:2128091633a81988_10_30]]
Based on the above data, answer the given subquestions.
[[IMAGE:2128091633a81988_12_35]]


A published solution is not available for this question yet.
Question 153 NAT · 1.0 marks
[[IMAGE:2128091633a81988_10_30]]
Based on the above data, answer the given subquestions.
How many deterministic optimal policies does this MDP have?

A published solution is not available for this question yet.
Question 154 NAT · 2.0 marks
[[IMAGE:2128091633a81988_14_36]]
Based on the above data, answer the given subquestions.
[[IMAGE:2128091633a81988_14_37]]


A published solution is not available for this question yet.
Question 155 NAT · 2.0 marks
[[IMAGE:2128091633a81988_14_36]]
Based on the above data, answer the given subquestions.
[[IMAGE:2128091633a81988_15_38]]


A published solution is not available for this question yet.