cs3004_2026T2_Q2_NA.pdf
Deep Learning · Quiz 2 · May 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 3.0 marks
[[IMAGE:897d6601db10d671_2_2]]

Bias will decrease significantly, while variance remains approximately the
same.
Variance will tend to decrease, while bias remains approximately the same.
Both bias and variance will eventually become zero.
Increasing the amount of training data has no effect on either the bias or the
variance.
A published solution is not available for this question yet.
Question 3 MCQ · 3.0 marks
[[IMAGE:897d6601db10d671_2_3]]

[[IMAGE:897d6601db10d671_3_4]] only

[[IMAGE:897d6601db10d671_3_5]]

[[IMAGE:897d6601db10d671_3_6]]

At all iterations
A published solution is not available for this question yet.
Question 4 NAT · 3.0 marks
[[IMAGE:897d6601db10d671_3_7]]
Based on the above data, answer the given subquestions.
Compute the effective learning rate [[IMAGE:897d6601db10d671_3_8]] used at
iteration [[IMAGE:897d6601db10d671_3_9]] . Enter the answer correct to four decimal places.



A published solution is not available for this question yet.
Question 5 MCQ · 3.0 marks
[[IMAGE:897d6601db10d671_3_7]]
Based on the above data, answer the given subquestions.
[[IMAGE:897d6601db10d671_4_10]]


It would increase from iteration 1 to iteration 3.
It would decrease from iteration 1 to iteration 3.
It would stay constant at [[IMAGE:897d6601db10d671_4_11]] .

It would decrease toward zero.
A published solution is not available for this question yet.
Question 6 MCQ · 3.0 marks
[[IMAGE:897d6601db10d671_4_12]]
Based on the above data, answer the given subquestions.
[[IMAGE:897d6601db10d671_4_13]]


[[IMAGE:897d6601db10d671_5_14]]

[[IMAGE:897d6601db10d671_5_15]]

[[IMAGE:897d6601db10d671_5_16]]

[[IMAGE:897d6601db10d671_5_17]]

A published solution is not available for this question yet.
Question 7 MCQ · 3.0 marks
[[IMAGE:897d6601db10d671_4_12]]
Based on the above data, answer the given subquestions.
[[IMAGE:897d6601db10d671_5_18]]


0
[[IMAGE:897d6601db10d671_5_19]]

[[IMAGE:897d6601db10d671_5_20]]

It diverges to [[IMAGE:897d6601db10d671_5_21]]

A published solution is not available for this question yet.
Question 8 NAT · 3.0 marks
A fully connected hidden layer has [[IMAGE:897d6601db10d671_5_22]] units and uses dropout with retention probability
[[IMAGE:897d6601db10d671_5_23]] for each unit (units are dropped independently). The sub-questions are independent of
each other.
Based on the above data, answer the given subquestions.
What is the expected number of units retained (active) in a single forward pass during training?


A published solution is not available for this question yet.
Question 9 MCQ · 3.0 marks
A fully connected hidden layer has [[IMAGE:897d6601db10d671_5_22]] units and uses dropout with retention probability
[[IMAGE:897d6601db10d671_5_23]] for each unit (units are dropped independently). The sub-questions are independent of
each other.
Based on the above data, answer the given subquestions.
How many distinct ''thinned'' sub-networks are theoretically possible for this layer?


[[IMAGE:897d6601db10d671_6_24]]

[[IMAGE:897d6601db10d671_6_25]]

[[IMAGE:897d6601db10d671_6_26]]

[[IMAGE:897d6601db10d671_6_27]]

A published solution is not available for this question yet.
Question 10 NAT · 2.0 marks
[[IMAGE:897d6601db10d671_6_28]]
Compute the pre-activation value [[IMAGE:897d6601db10d671_7_29]] .


A published solution is not available for this question yet.
Question 11 NAT · 2.0 marks
[[IMAGE:897d6601db10d671_6_28]]
Given your value of [[IMAGE:897d6601db10d671_7_30]] from the previous part, compute [[IMAGE:897d6601db10d671_7_31]] .Enter the answer correct to one
decimal place.



A published solution is not available for this question yet.
Question 12 MCQ · 2.0 marks
[[IMAGE:897d6601db10d671_6_28]]
Which of the following is the primary motivation for using Leaky ReLU instead of standard ReLU in
this setting?

To make the activation function computationally cheaper than ReLU.
To bound the output strictly within [[IMAGE:897d6601db10d671_7_32]] .

To allow a small, non-zero gradient to flow through when [[IMAGE:897d6601db10d671_7_33]] , avoiding
permanently dead neurons.

To make the activation zero-centered exactly like [[IMAGE:897d6601db10d671_7_34]] .

A published solution is not available for this question yet.
Question 13 MCQ · 2.0 marks
[[IMAGE:897d6601db10d671_8_35]]
For a training example where [[IMAGE:897d6601db10d671_8_36]] , which of the following best describes [[IMAGE:897d6601db10d671_8_37]] ?



[[IMAGE:897d6601db10d671_8_38]] ; the largest value [[IMAGE:897d6601db10d671_8_39]] can attain


[[IMAGE:897d6601db10d671_8_40]] ; the neuron has saturated

[[IMAGE:897d6601db10d671_8_41]]

[[IMAGE:897d6601db10d671_8_42]] is undefined at this point

A published solution is not available for this question yet.
Question 14 MCQ · 2.0 marks
[[IMAGE:897d6601db10d671_8_35]]
Suppose instead the same layer used ReLU, [[IMAGE:897d6601db10d671_8_43]] , and [[IMAGE:897d6601db10d671_8_44]] for most training
examples but a large negative bias update later drives [[IMAGE:897d6601db10d671_8_45]] for all future inputs. What is [[IMAGE:897d6601db10d671_8_46]] in
that regime, and what training consequence follows?





[[IMAGE:897d6601db10d671_8_47]] ; the neuron continues to update normally

[[IMAGE:897d6601db10d671_8_48]] ; the neuron stops updating and may stay ''dead'' permanently

[[IMAGE:897d6601db10d671_9_49]] oscillates between 0 and 1 depending on the input

[[IMAGE:897d6601db10d671_9_50]] ; the neuron updates at half the normal rate

A published solution is not available for this question yet.
Question 15 MSQ · 2.0 marks
[[IMAGE:897d6601db10d671_9_51]]
Based on the above data, answer the given subquestions.
Which filter weights can receive a non-zero gradient during backpropagation? (Select all that
apply.)

[[IMAGE:897d6601db10d671_9_52]]

[[IMAGE:897d6601db10d671_9_53]]

[[IMAGE:897d6601db10d671_9_54]]

[[IMAGE:897d6601db10d671_9_55]]

A published solution is not available for this question yet.
Question 16 NAT · 3.0 marks
[[IMAGE:897d6601db10d671_9_51]]
Based on the above data, answer the given subquestions.
[[IMAGE:897d6601db10d671_10_56]]
Suppose at each of the four output cells, and the current weights satisfy [[IMAGE:897d6601db10d671_10_57]] .
[[IMAGE:897d6601db10d671_10_58]]
Compute .




A published solution is not available for this question yet.
Question 17 NAT · 3.0 marks
The input volume to a convolutional layer of a CNN has shape [[IMAGE:897d6601db10d671_10_59]] , where [[IMAGE:897d6601db10d671_10_60]] is the depth.
Based on the above data, answer the given subquestions.
You want a convolutional layer with a [[IMAGE:897d6601db10d671_10_61]] kernel and stride [[IMAGE:897d6601db10d671_10_62]] to produce an output whose
width and height equal those of the input. What should be the padding [[IMAGE:897d6601db10d671_10_63]] ?





A published solution is not available for this question yet.
Question 18 MCQ · 2.0 marks
The input volume to a convolutional layer of a CNN has shape [[IMAGE:897d6601db10d671_10_59]] , where [[IMAGE:897d6601db10d671_10_60]] is the depth.
Based on the above data, answer the given subquestions.
Suppose you instead convolve the [[IMAGE:897d6601db10d671_11_64]] input with a [[IMAGE:897d6601db10d671_11_65]] kernel using padding [[IMAGE:897d6601db10d671_11_66]] ,
stride [[IMAGE:897d6601db10d671_11_67]] , and [[IMAGE:897d6601db10d671_11_68]] filters. Which of these is the shape of the output volume?







[[IMAGE:897d6601db10d671_11_69]]

[[IMAGE:897d6601db10d671_11_70]]

[[IMAGE:897d6601db10d671_11_71]]

[[IMAGE:897d6601db10d671_11_72]]

[[IMAGE:897d6601db10d671_11_73]]

A published solution is not available for this question yet.
Question 19 NAT · 2.0 marks
The input volume to a convolutional layer of a CNN has shape [[IMAGE:897d6601db10d671_10_59]] , where [[IMAGE:897d6601db10d671_10_60]] is the depth.
Based on the above data, answer the given subquestions.
Find the total number of parameters (weights [[IMAGE:897d6601db10d671_11_74]] biases) associated with the convolutional layer of
previous question.



A published solution is not available for this question yet.
Question 20 MSQ · 4.0 marks
Which of the following statements are TRUE about the adaptive optimizers? Select all that apply.
In AdaGrad, the effective learning rate [[IMAGE:897d6601db10d671_12_75]] decays slowly for the dense
features.

In AdaGrad, the effective learning rate [[IMAGE:897d6601db10d671_12_76]] decays slowly for the sparse
features.

In RMSProp, the effective learning rate [[IMAGE:897d6601db10d671_12_77]] is guaranteed to be non-
increasing across iterations.

[[IMAGE:897d6601db10d671_12_78]]

[[IMAGE:897d6601db10d671_12_79]]

A published solution is not available for this question yet.