da5007_2025T1_Q1_NA.pdf
Reinforcement Learning · Quiz 1 · Jan 2025
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 189 NAT · 1.0 marks
[[IMAGE:ef0e60a36fc290d9_1_1]]

A published solution is not available for this question yet.
Question 191 MCQ · 0.0 marks
**Note**:
For numerical answer type questions, enter your answer correct upto two decimal places without
rounding up or off unless stated otherwise.
Instructions has been mentioned above.
This Instructions is just for a reference & not for an evaluation.
A published solution is not available for this question yet.
Question 192 MCQ · 2.0 marks
Consider following assertion reason pair:
**Assertion**: Reinforcement learning is a type of unsupervised learning algorithm as both don’t have
correct labels.
**Reason**: In unsupervised learning, a reward like quantity is not maximized.
Choose the correct option.
Assertion and Reason are both true and Reason is a correct explanation of
Assertion.
Assertion and Reason are both true and Reason is not a correct explanation of
Assertion.
Assertion is true but Reason is false.
Assertion is false but Reason is true.
A published solution is not available for this question yet.
Question 193 MCQ · 2.0 marks
[[IMAGE:ef0e60a36fc290d9_3_2]]

1, 2, 3
only 3
only 1
1, 2
1, 3
A published solution is not available for this question yet.
Question 194 MCQ · 2.0 marks
[[IMAGE:ef0e60a36fc290d9_4_3]]

[[IMAGE:ef0e60a36fc290d9_4_4]]

[[IMAGE:ef0e60a36fc290d9_4_5]]

[[IMAGE:ef0e60a36fc290d9_4_6]]

[[IMAGE:ef0e60a36fc290d9_4_7]]

A published solution is not available for this question yet.
Question 195 MCQ · 2.0 marks
For any given Finite MDP, are all of its optimal policies necessarily deterministic?
TRUE
FALSE
A published solution is not available for this question yet.
Question 196 MCQ · 2.0 marks
Does changing the order of updates (over the entire set of states in each iteration) impact the
convergence of Asynchronous Value Iteration?
Yes
No
It depends
A published solution is not available for this question yet.
Question 197 MCQ · 3.0 marks
[[IMAGE:ef0e60a36fc290d9_5_8]]

[[IMAGE:ef0e60a36fc290d9_5_9]]

[[IMAGE:ef0e60a36fc290d9_5_10]]

[[IMAGE:ef0e60a36fc290d9_5_11]]

[[IMAGE:ef0e60a36fc290d9_5_12]]

A published solution is not available for this question yet.
Question 198 MCQ · 3.0 marks
[[IMAGE:ef0e60a36fc290d9_5_13]]

Arm 1
Arm 2
Arm 3
Arm 4
A published solution is not available for this question yet.
Question 199 MSQ · 2.0 marks
Select the correct
First-Visit MC is unbiased
Every-Visit MC is unbiased
First-Visit MC is more sample-efficient compared to Every-Visit MC
Every-Visit MC is more sample-efficient compared to First-Visit MC
A published solution is not available for this question yet.
Question 200 MSQ · 2.0 marks
Which of the following methods use bootstrapping?
TD
MC
DP
A published solution is not available for this question yet.
Question 201 MSQ · 2.0 marks
Which of the following methods use sampling?
TD
MC
DP
A published solution is not available for this question yet.
Question 202 NAT · 4.0 marks
[[IMAGE:ef0e60a36fc290d9_6_14]]

A published solution is not available for this question yet.
Question 203 NAT · 3.0 marks
[[IMAGE:ef0e60a36fc290d9_7_15]]

A published solution is not available for this question yet.
Question 204 MCQ · 2.0 marks
[[IMAGE:ef0e60a36fc290d9_7_16]]
Based on the above data, answer the given subquestions.
Assume **M = 10000**, then after the 10000 pull:

[[IMAGE:ef0e60a36fc290d9_8_17]]

[[IMAGE:ef0e60a36fc290d9_8_18]]

[[IMAGE:ef0e60a36fc290d9_8_19]]

[[IMAGE:ef0e60a36fc290d9_8_20]]

A published solution is not available for this question yet.
Question 205 MCQ · 2.0 marks
[[IMAGE:ef0e60a36fc290d9_7_16]]
Based on the above data, answer the given subquestions.
Assume **M = 10**, then after the 10000 pull:

[[IMAGE:ef0e60a36fc290d9_8_21]]

[[IMAGE:ef0e60a36fc290d9_8_22]]

[[IMAGE:ef0e60a36fc290d9_8_23]]

[[IMAGE:ef0e60a36fc290d9_8_24]]

A published solution is not available for this question yet.
Question 206 MCQ · 2.0 marks
[[IMAGE:ef0e60a36fc290d9_9_25]]
Based on the above data, answer the given subquestions.
Does there exist a stochastic (i.e. not deterministic) policy that is optimal?

TRUE
FALSE
A published solution is not available for this question yet.
Question 207 NAT · 3.0 marks
[[IMAGE:ef0e60a36fc290d9_9_25]]
Based on the above data, answer the given subquestions.
[[IMAGE:ef0e60a36fc290d9_9_26]]


A published solution is not available for this question yet.
Question 208 NAT · 3.0 marks
[[IMAGE:ef0e60a36fc290d9_10_27]]
[[IMAGE:ef0e60a36fc290d9_10_28]]


A published solution is not available for this question yet.
Question 209 NAT · 3.0 marks
[[IMAGE:ef0e60a36fc290d9_10_27]]
[[IMAGE:ef0e60a36fc290d9_10_29]]


A published solution is not available for this question yet.
Question 210 NAT · 3.0 marks
[[IMAGE:ef0e60a36fc290d9_11_30]]
Based on the above data, answer the given subquestions.
[[IMAGE:ef0e60a36fc290d9_11_31]]


A published solution is not available for this question yet.
Question 211 NAT · 3.0 marks
[[IMAGE:ef0e60a36fc290d9_11_30]]
Based on the above data, answer the given subquestions.
[[IMAGE:ef0e60a36fc290d9_11_32]]


A published solution is not available for this question yet.