cs3004_2026T1_Q2_NA.pdf
Deep Learning · Quiz 2 · Jan 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 1.0 marks
Which of the following optimization algorithms combines Nesterov accelerated gradient with
adaptive moment estimation?
NAdam
Adam
RMSprop with momentum
Nesterov accelerated gradient descent
A published solution is not available for this question yet.
Question 3 MCQ · 2.0 marks
Suppose that we apply Dropout regularization to a feed forward neural network. Suppose further
that mini-batch gradient descent algorithm is used for updating the parameters of the network.
Choose the correct statement from the following:
This will lead to sparcity in the trained weights.
The dropout probability [[IMAGE:6491eda98b90e1b2_2_2]] can be different for each hidden layer.

The weights of the neurons which were dropped during the forward
propagation at [[IMAGE:6491eda98b90e1b2_2_3]] -th iteration will not get updated during [[IMAGE:6491eda98b90e1b2_2_4]] -th iteration.


None of these.
A published solution is not available for this question yet.
Question 4 MCQ · 4.0 marks
Match the optimization algorithms in **Column A** with their corresponding properties or
characteristics in **Column B**. Note that an algorithm in Column A may correspond to multiple
properties in Column B.
[[IMAGE:6491eda98b90e1b2_3_5]]
Choose the correct set of mappings:

1-(i), 2-(v), 3-(ii, iii), 4-(iv)
1-(ii), 2-(iv), 3-(iii), 4-(i, v)
1-(ii), 2-(iii), 3-(i, iv), 4-(v)
1-(v), 2-(i), 3-(iv), 4-(ii, iii)
A published solution is not available for this question yet.
Question 5 MSQ · 2.0 marks
Consider training a deep neural network for a multi-class classification problem. The output layer
uses softmax activation, and the hidden layers use ReLU as an activation function. After a few
training iterations, you observe that the weights in the initial hidden layers barely get updated.
Which of the following actions could help tackle this issue?
Replace ReLU with Leaky ReLU in the hidden layers.
Replace ReLU by tanh in the hidden layers.
Increasing the number of neurons in the output layer.
Increasing the learning rate.
A published solution is not available for this question yet.
Question 6 MSQ · 4.0 marks
Consider a simple neural network with a hidden unit with ReLU as an activation function and a
linear output unit. The loss used is the squared loss: [[IMAGE:6491eda98b90e1b2_4_6]] .
[[IMAGE:6491eda98b90e1b2_4_7]]
Let [[IMAGE:6491eda98b90e1b2_4_8]] and [[IMAGE:6491eda98b90e1b2_4_9]] be the activations of the hidden and the output neurons, respectively.
[[IMAGE:6491eda98b90e1b2_4_10]] and [[IMAGE:6491eda98b90e1b2_4_11]] are the pre-activations for these neurons. Suppose we use
stochastic gradient descent to update the weights and biases with a learning rate [[IMAGE:6491eda98b90e1b2_4_12]] . Then,
which among the following is true?







[[IMAGE:6491eda98b90e1b2_4_13]]

[[IMAGE:6491eda98b90e1b2_4_14]]

[[IMAGE:6491eda98b90e1b2_4_15]] will remain unchanged after one SGD update if [[IMAGE:6491eda98b90e1b2_4_16]] .


[[IMAGE:6491eda98b90e1b2_4_17]] will remain unchanged after one SGD update if [[IMAGE:6491eda98b90e1b2_4_18]] .


[[IMAGE:6491eda98b90e1b2_4_19]] will remain unchanged after one SGD update if [[IMAGE:6491eda98b90e1b2_4_20]] and at the current
step [[IMAGE:6491eda98b90e1b2_4_21]] .



[[IMAGE:6491eda98b90e1b2_5_22]] will remain unchanged after one SGD update if [[IMAGE:6491eda98b90e1b2_5_23]] .


A published solution is not available for this question yet.
Question 7 NAT · 3.0 marks
Consider a convolutional neural network architecture where an input image of size
[[IMAGE:6491eda98b90e1b2_5_24]] is passed into two separate, parallel convolutional layers ( [[IMAGE:6491eda98b90e1b2_5_25]] and [[IMAGE:6491eda98b90e1b2_5_26]] )
simultaneously. Both layers are configured with the following specifications:
• **Number of Filters:** 64
• **Kernel Size:** [[IMAGE:6491eda98b90e1b2_5_27]]
• **Stride:** 1
• **Bias:** Assume no bias parameters are used.
The only difference between the two layers is the padding configuration:
• **Layer** [[IMAGE:6491eda98b90e1b2_5_28]] : Uses Zero Padding ( [[IMAGE:6491eda98b90e1b2_5_29]] ).
• **Layer** [[IMAGE:6491eda98b90e1b2_5_30]] : Uses Padding of 2 ( [[IMAGE:6491eda98b90e1b2_5_31]] ).
Calculate the difference in the total number of learnable parameters between Layer [[IMAGE:6491eda98b90e1b2_5_32]] and Layer
[[IMAGE:6491eda98b90e1b2_5_33]] .










A published solution is not available for this question yet.
Question 8 NAT · 3.0 marks
[[IMAGE:6491eda98b90e1b2_6_34]]

A published solution is not available for this question yet.
Question 9 NAT · 2.0 marks
[[IMAGE:6491eda98b90e1b2_6_35]]
.
Based on the above data, answer the given subquestions.
Compute Bias [[IMAGE:6491eda98b90e1b2_7_36]] .


A published solution is not available for this question yet.
Question 10 NAT · 2.0 marks
[[IMAGE:6491eda98b90e1b2_6_35]]
.
Based on the above data, answer the given subquestions.
Compute Variance [[IMAGE:6491eda98b90e1b2_7_37]] . Enter the answer correct to four decimal places.


A published solution is not available for this question yet.
Question 11 NAT · 4.0 marks
Consider a neural network with one hidden layer as shown below:
[[IMAGE:6491eda98b90e1b2_8_38]]
The input layer consist of 3 features [[IMAGE:6491eda98b90e1b2_8_39]] , the hidden layer consist of two neurons
and the output layer consist of 3 neurons. Following are the weights and biases of the network:
[[IMAGE:6491eda98b90e1b2_8_40]]
[[IMAGE:6491eda98b90e1b2_8_41]]
[[IMAGE:6491eda98b90e1b2_8_42]]
.
Based on the above data, answer the given subquestions.
Do a forward propagation through the network for a training example [[IMAGE:6491eda98b90e1b2_9_43]] and
[[IMAGE:6491eda98b90e1b2_9_44]] and compute the loss. Use natural log for the calculation. Enter the answer correct
to two decimal places.







A published solution is not available for this question yet.
Question 12 NAT · 3.0 marks
Consider a neural network with one hidden layer as shown below:
[[IMAGE:6491eda98b90e1b2_8_38]]
The input layer consist of 3 features [[IMAGE:6491eda98b90e1b2_8_39]] , the hidden layer consist of two neurons
and the output layer consist of 3 neurons. Following are the weights and biases of the network:
[[IMAGE:6491eda98b90e1b2_8_40]]
[[IMAGE:6491eda98b90e1b2_8_41]]
[[IMAGE:6491eda98b90e1b2_8_42]]
.
Based on the above data, answer the given subquestions.
Suppose now we introduce an L2 regularization into the loss function. The regularised loss will be
[[IMAGE:6491eda98b90e1b2_9_45]]
where [[IMAGE:6491eda98b90e1b2_9_46]] is a regularization parameter, and [[IMAGE:6491eda98b90e1b2_9_47]] is the weight of the network where
[[IMAGE:6491eda98b90e1b2_9_48]] Compute the regularised loss for the training example [[IMAGE:6491eda98b90e1b2_9_49]]
assuming [[IMAGE:6491eda98b90e1b2_9_50]] . Use weights before the backpropagation.











A published solution is not available for this question yet.
Question 13 MSQ · 3.0 marks
Consider a neural network with one hidden layer as shown below:
[[IMAGE:6491eda98b90e1b2_8_38]]
The input layer consist of 3 features [[IMAGE:6491eda98b90e1b2_8_39]] , the hidden layer consist of two neurons
and the output layer consist of 3 neurons. Following are the weights and biases of the network:
[[IMAGE:6491eda98b90e1b2_8_40]]
[[IMAGE:6491eda98b90e1b2_8_41]]
[[IMAGE:6491eda98b90e1b2_8_42]]
.
Based on the above data, answer the given subquestions.
Based on the previous questions 11 and 12, select the correct options from the following:





[[IMAGE:6491eda98b90e1b2_9_51]]

Gradient updates will be larger in case of regularised loss.
L [[IMAGE:6491eda98b90e1b2_9_52]] regularisation introduces weight decay during SGD updates.

As [[IMAGE:6491eda98b90e1b2_10_53]] will increase.

A published solution is not available for this question yet.
Question 14 NAT · 2.0 marks
Consider two neural network modules, **Module A** and **Module B**, which are designed to process
an input feature map of size [[IMAGE:6491eda98b90e1b2_10_54]] .
• **Module A** utilizes a serial bottleneck design.
• **Module B** utilizes a parallel Inception-style design.
**Note:** For all questions, assume stride = 1, "same'' padding, and no bias parameters.
[[IMAGE:6491eda98b90e1b2_10_55]]
Based on the above data, answer the given subquestions.
Calculate the total number of learnable parameters in **Module A**.


A published solution is not available for this question yet.
Question 15 NAT · 2.0 marks
Consider two neural network modules, **Module A** and **Module B**, which are designed to process
an input feature map of size [[IMAGE:6491eda98b90e1b2_10_54]] .
• **Module A** utilizes a serial bottleneck design.
• **Module B** utilizes a parallel Inception-style design.
**Note:** For all questions, assume stride = 1, "same'' padding, and no bias parameters.
[[IMAGE:6491eda98b90e1b2_10_55]]
Based on the above data, answer the given subquestions.
Calculate the total number of learnable parameters in **Module B**.


A published solution is not available for this question yet.
Question 16 MCQ · 3.0 marks
Consider two neural network modules, **Module A** and **Module B**, which are designed to process
an input feature map of size [[IMAGE:6491eda98b90e1b2_10_54]] .
• **Module A** utilizes a serial bottleneck design.
• **Module B** utilizes a parallel Inception-style design.
**Note:** For all questions, assume stride = 1, "same'' padding, and no bias parameters.
[[IMAGE:6491eda98b90e1b2_10_55]]
Based on the above data, answer the given subquestions.
What is the primary functional advantage of the [[IMAGE:6491eda98b90e1b2_11_56]] convolution in **Module A** compared to its
role in a standard Inception module?



It increases the spatial resolution of the feature map.
It acts as a bottleneck layer to reduce the computational cost of the
subsequent [[IMAGE:6491eda98b90e1b2_11_57]] convolution.

It allows the network to learn parallel features at different scales.
It is only used to add non-linearity and does not affect the parameter count.
A published solution is not available for this question yet.