cs3004_2025T3_Q2_NA.pdf
Deep Learning · Quiz 2 · Sep 2025
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 71 MSQ · 1.0 marks
Which of the following are **relevant** actions for the given planning problem?
Pickup(A)
Putdown(A)
Putdown(B)
Putdown(C)
Stack(A,B)
A published solution is not available for this question yet.
Question 72 MCQ · 1.0 marks
Which of the following can be pushed as the first three elements onto the stack by the Goal Stack
Planning algorithm? In the stack representation, the bottom is on the right side marked by the
entry BOTTOM. Use appropriate tie-breakers listed in the main data.
{ onTable(B), on(A,B), onTable(C) }; onTable(B); on(A,B); onTable(C); BOTTOM
{ onTable(B), on(A,B), onTable(C) }; onTable(C); on(A,B); onTable(B); BOTTOM
onTable(B); on(A,B); onTable(C); { onTable(B), on(A,B), onTable(C) }; BOTTOM
onTable(C); on(A,B); onTable(B); { onTable(B), on(A,B), onTable(C) }; BOTTOM
A published solution is not available for this question yet.
Question 73 MCQ · 1.0 marks
For the subgoal ordering given in the goal description (and using the given tie breaking rules),
which of the following is the first action popped out of the stack in Goal Stack Planning?
Pickup(A)
Putdown(A)
Putdown(B)
Putdown(C)
Stack(A,B)
**Deep Learning**
**Section Id :** 640653121909
**Section Number :** 4
**Section type :** Online
**Mandatory or Optional :** Mandatory
**Number of Questions :** 11
**Number of Questions to be attempted :** 11
**Section Marks :** 40
**Display Number Panel :** Yes
**Section Negative Marks :** 0
**Group All Questions :** No
**Enable Mark as Answered Mark for Review and**
No
**Clear Response :**
**Section Maximum Duration :** 0
**Section Minimum Duration :** 0
**Section Time In :** Minutes
**Maximum Instruction Time :** 0
A published solution is not available for this question yet.
Question 75 MCQ · 3.0 marks
[[IMAGE:6e81992fcfcaec04_3_0]]

Choose [[IMAGE:6e81992fcfcaec04_3_1]] , since it gives the lowest training loss and allows the model
to fit the data best.

Choose [[IMAGE:6e81992fcfcaec04_3_2]] , as it provides a reasonable compromise between model
complexity and generalization.

Choose [[IMAGE:6e81992fcfcaec04_3_3]] , since it yields the lowest validation loss and achieves the best
bias–variance trade-off.

Choose [[IMAGE:6e81992fcfcaec04_3_4]] , as it produces a highly regularized model that minimizes the
risk of overfitting.

A published solution is not available for this question yet.
Question 76 MCQ · 4.0 marks
Consider a neural network as shown in the figure with 2 input neurons, a hidden layer with 2
neurons (ReLU), and 1 output neuron (sigmoid).
Weights are:
[[IMAGE:6e81992fcfcaec04_3_5]]
with no biases. Dropout probability is 0.5 on the hidden layer.
If the input vector [[IMAGE:6e81992fcfcaec04_3_6]] is fed to this network, what is the expected output while testing?
[[IMAGE:6e81992fcfcaec04_4_7]]



0.77
1.25
2.5
0.92
A published solution is not available for this question yet.
Question 77 MSQ · 3.0 marks
Which of the following statements are true about a Convolutional (CONV) layer?
The number of parameters depends on the depth of the input.
The number of filters determines the number of output channels.
The total number of parameters depends on the stride.
The total number of parameters depends on the padding.
A published solution is not available for this question yet.
Question 78 MSQ · 3.0 marks
Which of the following optimizers adapts the learning rate for each parameter during training?
Stochastic Gradient Descent
Nesterov Accelerated Gradient Descent
Adagrad
RMSProp
A published solution is not available for this question yet.
Question 79 MSQ · 3.0 marks
Which of the following techniques can be used to reduce model overfitting?
Data augmentation
Dropout
L2 regularization
Using Adam instead of SGD
A published solution is not available for this question yet.
Question 80 MCQ · 3.0 marks
Why is bias correction needed in the Adam optimizer?
To prevent the learning rate from becoming too large in the early iterations.
To correct the underestimated moving averages of the first and second
moments at the beginning of training.
To increase the momentum effect for faster convergence.
To ensure the learning rate remains constant throughout training.
A published solution is not available for this question yet.
Question 81 MSQ · 3.0 marks
After training a linear regression model on a large training set of size [[IMAGE:6e81992fcfcaec04_6_8]] , it achieves a training
error of [[IMAGE:6e81992fcfcaec04_6_9]] . Analysis of the residual plot shows a clear non-linear pattern, suggesting the model is
underfitting the data. Which two of the following modifications are most likely to improve the
model’s performance by increasing its capacity to capture non-linear relationships?


Adding a regularization term (such as L2 or L1 penalty) to the mean squared
error loss function.
Using a 2-hidden layer feedforward network with ReLU activation functions in
place of linear regression.
Applying a polynomial feature transformation of degree [[IMAGE:6e81992fcfcaec04_6_10]] to the input
variables.

Standardizing all training samples to have mean zero and unit variance.
Using a 5-hidden layer feedforward network without non-linear activation
functions in place of linear regression.
A published solution is not available for this question yet.
Question 82 NAT · 4.0 marks
Consider an Inception module in GoogLeNet with an input feature map of size [[IMAGE:6e81992fcfcaec04_6_11]] .
The module has four branches:
• Branch 1: [[IMAGE:6e81992fcfcaec04_6_12]] convolution with 64 filters.
• Branch 2: [[IMAGE:6e81992fcfcaec04_6_13]] convolution with 96 filters followed by [[IMAGE:6e81992fcfcaec04_6_14]] convolution with 128 filters.
• Branch 3: [[IMAGE:6e81992fcfcaec04_6_15]] convolution with 16 filters followed by [[IMAGE:6e81992fcfcaec04_6_16]] convolution with 32 filters.
• Branch 4: [[IMAGE:6e81992fcfcaec04_6_17]] max pooling followed by [[IMAGE:6e81992fcfcaec04_6_18]] convolution with 32 filters.
Assume that all convolutions use no padding, stride 1, and that biases are ignored. What is the
total number of learnable parameters in this module?








A published solution is not available for this question yet.
Question 83 MSQ · 2.0 marks
Consider a fully connected neural network as follows:
• One input neuron [[IMAGE:6e81992fcfcaec04_7_19]]
• A single hidden layer with two neurons
• One output neuron.
The network uses the ReLU activation function in both the hidden and output layers. The loss
function used is the squared error:
[[IMAGE:6e81992fcfcaec04_7_20]]
where [[IMAGE:6e81992fcfcaec04_7_21]] is the true target and [[IMAGE:6e81992fcfcaec04_7_22]] is the network output.
[[IMAGE:6e81992fcfcaec04_7_23]]
Let the initial weights of the network be [[IMAGE:6e81992fcfcaec04_7_24]] .
Based on the above data, answer the given subquestions.
Choose the option which can represent the loss for single input data.






[[IMAGE:6e81992fcfcaec04_8_25]] , if [[IMAGE:6e81992fcfcaec04_8_26]]


[[IMAGE:6e81992fcfcaec04_8_27]] , if [[IMAGE:6e81992fcfcaec04_8_28]]


[[IMAGE:6e81992fcfcaec04_8_29]] , if [[IMAGE:6e81992fcfcaec04_8_30]]


None of these
A published solution is not available for this question yet.
Question 84 MCQ · 2.0 marks
Consider a fully connected neural network as follows:
• One input neuron [[IMAGE:6e81992fcfcaec04_7_19]]
• A single hidden layer with two neurons
• One output neuron.
The network uses the ReLU activation function in both the hidden and output layers. The loss
function used is the squared error:
[[IMAGE:6e81992fcfcaec04_7_20]]
where [[IMAGE:6e81992fcfcaec04_7_21]] is the true target and [[IMAGE:6e81992fcfcaec04_7_22]] is the network output.
[[IMAGE:6e81992fcfcaec04_7_23]]
Let the initial weights of the network be [[IMAGE:6e81992fcfcaec04_7_24]] .
Based on the above data, answer the given subquestions.
Select the gradient of the loss function with respect to weight [[IMAGE:6e81992fcfcaec04_8_31]] from the following.
**Note:** [[IMAGE:6e81992fcfcaec04_8_32]] represents the indicator function.








[[IMAGE:6e81992fcfcaec04_8_33]]

[[IMAGE:6e81992fcfcaec04_8_34]]

[[IMAGE:6e81992fcfcaec04_8_35]]

[[IMAGE:6e81992fcfcaec04_8_36]]

A published solution is not available for this question yet.
Question 85 NAT · 3.0 marks
Consider a fully connected neural network as follows:
• One input neuron [[IMAGE:6e81992fcfcaec04_7_19]]
• A single hidden layer with two neurons
• One output neuron.
The network uses the ReLU activation function in both the hidden and output layers. The loss
function used is the squared error:
[[IMAGE:6e81992fcfcaec04_7_20]]
where [[IMAGE:6e81992fcfcaec04_7_21]] is the true target and [[IMAGE:6e81992fcfcaec04_7_22]] is the network output.
[[IMAGE:6e81992fcfcaec04_7_23]]
Let the initial weights of the network be [[IMAGE:6e81992fcfcaec04_7_24]] .
Based on the above data, answer the given subquestions.
Suppose the network uses the ELU activation function with [[IMAGE:6e81992fcfcaec04_8_37]] in the hidden layer and ReLU
activation function in the output layer. Using the same initial weights, compute the updated value
of [[IMAGE:6e81992fcfcaec04_8_38]] after one step of stochastic gradient descent on the data point [[IMAGE:6e81992fcfcaec04_8_39]] with a
learning rate of [[IMAGE:6e81992fcfcaec04_8_40]] . Enter the answer correct to one decimal place.
**Hint:** ELU activation function is given by
[[IMAGE:6e81992fcfcaec04_9_41]]











A published solution is not available for this question yet.
Question 86 MCQ · 2.0 marks
[[IMAGE:6e81992fcfcaec04_10_42]]
Based on the above data,answer the given subquestion.
What will be the **input** to the **Pooling Layer**.

[[IMAGE:6e81992fcfcaec04_10_43]]

[[IMAGE:6e81992fcfcaec04_11_44]]

[[IMAGE:6e81992fcfcaec04_11_45]]

[[IMAGE:6e81992fcfcaec04_11_46]]

A published solution is not available for this question yet.
Question 87 NAT · 2.0 marks
[[IMAGE:6e81992fcfcaec04_10_42]]
Based on the above data,answer the given subquestion.
Find the output after average pooling. Answer correct up to 2 digits after the decimal.

A published solution is not available for this question yet.
Question 88 NAT · 1.0 marks
[[IMAGE:6e81992fcfcaec04_10_42]]
Based on the above data,answer the given subquestion.
If [[IMAGE:6e81992fcfcaec04_11_47]] , compute [[IMAGE:6e81992fcfcaec04_11_48]] and submit the answer correct up to 2 digits after the decimal



A published solution is not available for this question yet.