cs3004_2025T3_Q1_NA.pdf
Deep Learning · Quiz 1 · Sep 2025
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 43 MSQ · 0.0 marks
Do you love this site 😍?
Yes
Option A
A published solution is not available for this question yet.
Question 45 MCQ · 4.0 marks
[[IMAGE:2e5f813e2ea0a2ee_2_2]]

[[IMAGE:2e5f813e2ea0a2ee_2_3]]

[[IMAGE:2e5f813e2ea0a2ee_2_4]]

[[IMAGE:2e5f813e2ea0a2ee_2_5]]

[[IMAGE:2e5f813e2ea0a2ee_3_6]]

[[IMAGE:2e5f813e2ea0a2ee_3_7]]

A published solution is not available for this question yet.
Question 46 MCQ · 3.0 marks
[[IMAGE:2e5f813e2ea0a2ee_3_8]]

A → 0.01 , B → 0.001 , C → 1.0
A → 0.001 , B → 0.01 , C → 1.0
A → 1.0 , B → 0.001 , C → 0.01
A → 0.01 , B → 1.0 , C → 0.001
A published solution is not available for this question yet.
Question 47 MCQ · 3.0 marks
Which of the following statements best describes the primary representational advantage of a
Multi-Layer Perceptron (MLP) over a Single-Layer Perceptron (SLP)?
An MLP can learn non-linear decision boundaries, allowing it to classify data
that is not linearly separable.
An MLP converges faster than an SLP on linearly separable data because its
hidden layers accelerate learning.
An MLP is less prone to overfitting than an SLP because it has more
parameters to capture the data distribution.
An MLP can only be used for classification tasks, whereas an SLP can be used
for both classification and regression.
A published solution is not available for this question yet.
Question 48 MCQ · 3.0 marks
A data scientist is training a deep learning model on a massive dataset containing billions of data
points. They find that using traditional Batch Gradient Descent (GD) is computationally infeasible.
What is the primary reason that makes Batch GD impractical for this task, thus requiring an
alternative like Stochastic Gradient Descent (SGD)?
It is impossible to fit the entire large dataset into the computer’s memory
(RAM) at once.
The gradient calculated by Batch GD is often too noisy and inaccurate, leading
to poor convergence.
Batch GD is guaranteed to get stuck in sharp local minima, whereas SGD can
escape them.
Batch GD requires calculating the gradient of the loss function with respect to
the parameters over the entire dataset for a single update, which is computationally very
expensive.
A published solution is not available for this question yet.
Question 49 MCQ · 3.0 marks
[[IMAGE:2e5f813e2ea0a2ee_4_9]]

[[IMAGE:2e5f813e2ea0a2ee_5_10]]

[[IMAGE:2e5f813e2ea0a2ee_5_11]]

[[IMAGE:2e5f813e2ea0a2ee_5_12]]

[[IMAGE:2e5f813e2ea0a2ee_5_13]]

A published solution is not available for this question yet.
Question 50 MSQ · 3.0 marks
You are adapting a neural network that was originally designed for a 10-class image classification
problem to perform a regression task, specifically to predict the price of a car. Which of the
following modifications are essential or standard practice for this conversion?
Change the activation function of the final output layer to a linear function.
Set the number of neurons in the output layer to one.
Change the loss function from Categorical Cross-Entropy to Mean Squared
Error (MSE).
Use accuracy as the primary metric to evaluate the model’s performance.
Keep the softmax activation function in the output layer to scale the predicted
price.
A published solution is not available for this question yet.
Question 51 NAT · 3.0 marks
[[IMAGE:2e5f813e2ea0a2ee_6_14]]

A published solution is not available for this question yet.
Question 52 NAT · 4.0 marks
[[IMAGE:2e5f813e2ea0a2ee_6_15]]

A published solution is not available for this question yet.
Question 53 NAT · 1.0 marks
[[IMAGE:2e5f813e2ea0a2ee_8_16]]
Based on the above data, answer the given subquestions.
What is the value of the output of the hidden layer? (Answer correct upto two digits after the
decimal)

A published solution is not available for this question yet.
Question 54 NAT · 1.0 marks
[[IMAGE:2e5f813e2ea0a2ee_8_16]]
Based on the above data, answer the given subquestions.
What is the value of the cross entropy loss? (use natural log). (Answer correct upto two digits after
the decimal)

A published solution is not available for this question yet.
Question 55 NAT · 2.0 marks
[[IMAGE:2e5f813e2ea0a2ee_8_16]]
Based on the above data, answer the given subquestions.
[[IMAGE:2e5f813e2ea0a2ee_9_17]]


A published solution is not available for this question yet.
Question 56 NAT · 2.0 marks
[[IMAGE:2e5f813e2ea0a2ee_8_16]]
Based on the above data, answer the given subquestions.
Use cross entropy loss and compute the gradient of W1 which is the weight between input and
hidden layer. (consider upto two digits after the decimal for all the calculations)

A published solution is not available for this question yet.
Question 57 MCQ · 1.0 marks
[[IMAGE:2e5f813e2ea0a2ee_10_18]]
Based on the above data, answer the given subquestions.
[[IMAGE:2e5f813e2ea0a2ee_10_19]]


[[IMAGE:2e5f813e2ea0a2ee_10_20]]

[[IMAGE:2e5f813e2ea0a2ee_11_21]]

[[IMAGE:2e5f813e2ea0a2ee_11_22]]

[[IMAGE:2e5f813e2ea0a2ee_11_23]]

A published solution is not available for this question yet.
Question 58 MCQ · 1.0 marks
[[IMAGE:2e5f813e2ea0a2ee_10_18]]
Based on the above data, answer the given subquestions.
[[IMAGE:2e5f813e2ea0a2ee_11_24]]


[[IMAGE:2e5f813e2ea0a2ee_11_25]]

[[IMAGE:2e5f813e2ea0a2ee_11_26]]

[[IMAGE:2e5f813e2ea0a2ee_11_27]]

[[IMAGE:2e5f813e2ea0a2ee_11_28]]

A published solution is not available for this question yet.
Question 59 MCQ · 1.0 marks
[[IMAGE:2e5f813e2ea0a2ee_10_18]]
Based on the above data, answer the given subquestions.
[[IMAGE:2e5f813e2ea0a2ee_11_29]]


[[IMAGE:2e5f813e2ea0a2ee_11_30]]

[[IMAGE:2e5f813e2ea0a2ee_11_31]]

[[IMAGE:2e5f813e2ea0a2ee_11_32]]

[[IMAGE:2e5f813e2ea0a2ee_11_33]]

A published solution is not available for this question yet.
Question 60 MCQ · 1.0 marks
[[IMAGE:2e5f813e2ea0a2ee_10_18]]
Based on the above data, answer the given subquestions.
[[IMAGE:2e5f813e2ea0a2ee_12_34]]


[[IMAGE:2e5f813e2ea0a2ee_12_35]]

[[IMAGE:2e5f813e2ea0a2ee_12_36]]

[[IMAGE:2e5f813e2ea0a2ee_12_37]]

[[IMAGE:2e5f813e2ea0a2ee_12_38]]

A published solution is not available for this question yet.
Question 61 MCQ · 2.0 marks
Suppose you build a neural network with one hidden layer that uses sigmoid as an activation
function and an output layer with a softmax activation function. You initialize the hidden layer
weights W1 randomly and set the output layer weights W2 to all zeros. Assume the bias is set to be
zero. You train the network using stochastic gradient descent.
Based on the above data, answer the given subquestions.
After one iteration of gradient descent, will the new weight for W1 be the same as the previous
weight?
Yes
No
A published solution is not available for this question yet.
Question 62 MCQ · 2.0 marks
Suppose you build a neural network with one hidden layer that uses sigmoid as an activation
function and an output layer with a softmax activation function. You initialize the hidden layer
weights W1 randomly and set the output layer weights W2 to all zeros. Assume the bias is set to be
zero. You train the network using stochastic gradient descent.
Based on the above data, answer the given subquestions.
After one iteration of gradient descent, will the new weight for W2 be the same as the previous
weight?
Yes
No
A published solution is not available for this question yet.