cs3004_2026T1_Q1_NA.pdf
Deep Learning · Quiz 1 · Jan 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 2.0 marks
Which among the following is true for a McCulloch-Pitts (MP) neuron?
It can implement Boolean functions with a non-linear decision boundary.
It can implement linearly separable Boolean functions with a linear decision
boundary.
It can be used to approximate a real-valued function.
It can accurately represent the XOR function by adjusting its weights and
thresholds.
A published solution is not available for this question yet.
Question 3 MCQ · 4.0 marks
Consider a shallow neural network with a scalar input [[IMAGE:8f61ac644f74e0fb_3_2]] and a hidden layer with three neurons.
The outputs of the hidden layer neurons are given by
[[IMAGE:8f61ac644f74e0fb_3_3]]
where [[IMAGE:8f61ac644f74e0fb_4_4]] denotes the activation function. The final output of the network is given by
[[IMAGE:8f61ac644f74e0fb_4_5]]
Suppose we multiply the parameters [[IMAGE:8f61ac644f74e0fb_4_6]] and [[IMAGE:8f61ac644f74e0fb_4_7]] by a constant [[IMAGE:8f61ac644f74e0fb_4_8]] , and divide [[IMAGE:8f61ac644f74e0fb_4_9]] by the
same constant [[IMAGE:8f61ac644f74e0fb_4_10]] , while all other parameters remain unchanged. Which of the following
statements is correct?
[[IMAGE:8f61ac644f74e0fb_4_11]]










The network output [[IMAGE:8f61ac644f74e0fb_4_12]] remains unchanged if the activation function used is
ReLU for [[IMAGE:8f61ac644f74e0fb_4_13]] , [[IMAGE:8f61ac644f74e0fb_4_14]] .



The network output [[IMAGE:8f61ac644f74e0fb_4_15]] remains unchanged if the activation function used is
ReLU for [[IMAGE:8f61ac644f74e0fb_4_16]] , [[IMAGE:8f61ac644f74e0fb_4_17]] .



The network output will be unchanged if the activation function used is
sigmoid for [[IMAGE:8f61ac644f74e0fb_4_18]] .

After the transformation, the contribution [[IMAGE:8f61ac644f74e0fb_4_19]] to the output remains
unchanged for all [[IMAGE:8f61ac644f74e0fb_4_20]] and for all [[IMAGE:8f61ac644f74e0fb_4_21]] .



A published solution is not available for this question yet.
Question 4 MSQ · 3.0 marks
Consider a multilayer perceptron (MLP) with two binary inputs [[IMAGE:8f61ac644f74e0fb_4_22]] and [[IMAGE:8f61ac644f74e0fb_4_23]] , 2 neurons in the hidden
layer, and 1 in the output layer. Let [[IMAGE:8f61ac644f74e0fb_4_24]] and [[IMAGE:8f61ac644f74e0fb_4_25]] be the outputs of the hidden layer and it computes
[[IMAGE:8f61ac644f74e0fb_4_26]]
The output of the network is given by [[IMAGE:8f61ac644f74e0fb_5_27]] , where
[[IMAGE:8f61ac644f74e0fb_5_28]]
Which of the following are the possible choices of [[IMAGE:8f61ac644f74e0fb_5_29]] for the network to correctly compute the
XOR function, given that [[IMAGE:8f61ac644f74e0fb_5_30]] .









0.5
-0.55
-0.60
1
A published solution is not available for this question yet.
Question 5 NAT · 2.0 marks
Consider a neural network for a regression problem with input [[IMAGE:8f61ac644f74e0fb_5_31]] . The network has [[IMAGE:8f61ac644f74e0fb_5_32]]
hidden layers, each with [[IMAGE:8f61ac644f74e0fb_5_33]] sigmoid neurons. The output layer has [[IMAGE:8f61ac644f74e0fb_5_34]] neuron and uses the sigmoid
activation function. Enter the number of parameters in the network.




A published solution is not available for this question yet.
Question 6 NAT · 3.0 marks
Consider a training set consisting of 12 samples. The vanilla gradient descent algorithm is used to
update the model parameters over 4 epochs. The learning rate follows an exponential decay
scheme given by
[[IMAGE:8f61ac644f74e0fb_6_35]]
where the iteration index [[IMAGE:8f61ac644f74e0fb_6_36]] starts from zero and increments only when the parameters are
updated. What is the value of the learning rate [[IMAGE:8f61ac644f74e0fb_6_37]] at the end of training?
(Note: Enter your answer up to 3 decimal points. For example, if your answer is 0.12345, enter 0.123).



A published solution is not available for this question yet.
Question 7 NAT · 4.0 marks
A content recommendation system models whether a user will engage with an item based on
features [[IMAGE:8f61ac644f74e0fb_6_38]] .
The training dataset is
[[IMAGE:8f61ac644f74e0fb_6_39]]
where [[IMAGE:8f61ac644f74e0fb_6_40]] denotes non-engagement (0) or engagement (1). The probability of
engagement is modeled using:
[[IMAGE:8f61ac644f74e0fb_6_41]]
where [[IMAGE:8f61ac644f74e0fb_6_42]] , and [[IMAGE:8f61ac644f74e0fb_6_43]] denote the bias. The model parameters are learned using gradient descent
by minimizing the squared error loss:
[[IMAGE:8f61ac644f74e0fb_6_44]]
[[IMAGE:8f61ac644f74e0fb_6_45]]
Consider a training example that takes in an input vector , with true label [[IMAGE:8f61ac644f74e0fb_6_46]] . The
[[IMAGE:8f61ac644f74e0fb_6_48]]
weight vector [[IMAGE:8f61ac644f74e0fb_6_47]] is initialized to and [[IMAGE:8f61ac644f74e0fb_6_49]] . Compute the loss after one update of [[IMAGE:8f61ac644f74e0fb_6_50]] and [[IMAGE:8f61ac644f74e0fb_6_51]] for
[[IMAGE:8f61ac644f74e0fb_7_52]] . Enter the answer correct to two decimal places.















A published solution is not available for this question yet.
Question 8 MCQ · 1.0 marks
Consider the truth table of a NAND function for two binary inputs:
[[IMAGE:8f61ac644f74e0fb_7_53]]
You use a perceptron for implementing the NAND function. The output of the neuron is given as
[[IMAGE:8f61ac644f74e0fb_7_54]]
[[IMAGE:8f61ac644f74e0fb_7_55]] is assumed to be 1.
Based on the above data, answer the given subquestions.
Is the function linearly separable?



Yes
No
A published solution is not available for this question yet.
Question 9 MSQ · 3.0 marks
Consider the truth table of a NAND function for two binary inputs:
[[IMAGE:8f61ac644f74e0fb_7_53]]
You use a perceptron for implementing the NAND function. The output of the neuron is given as
[[IMAGE:8f61ac644f74e0fb_7_54]]
[[IMAGE:8f61ac644f74e0fb_7_55]] is assumed to be 1.
Based on the above data, answer the given subquestions.
Which of the following choices for the weights [[IMAGE:8f61ac644f74e0fb_8_56]] , [[IMAGE:8f61ac644f74e0fb_8_57]] , and [[IMAGE:8f61ac644f74e0fb_8_58]] will produce the NAND function?






[[IMAGE:8f61ac644f74e0fb_8_59]]

[[IMAGE:8f61ac644f74e0fb_8_60]]

[[IMAGE:8f61ac644f74e0fb_8_61]]

[[IMAGE:8f61ac644f74e0fb_8_62]]

A published solution is not available for this question yet.
Question 10 MCQ · 3.0 marks
Suppose you perform the perceptron algorithm on the following dataset:
[[IMAGE:8f61ac644f74e0fb_8_63]]
Assume the initial weights [[IMAGE:8f61ac644f74e0fb_8_64]] and we pass the data points in the order given in the
table. The following rule is used for the classification:
[[IMAGE:8f61ac644f74e0fb_8_65]]
Based on the above data, answer the given subquestions.
What will be the updated weight vector after one epoch?



[[IMAGE:8f61ac644f74e0fb_9_66]]

[[IMAGE:8f61ac644f74e0fb_9_67]]

[[IMAGE:8f61ac644f74e0fb_9_68]]

[[IMAGE:8f61ac644f74e0fb_9_69]]

A published solution is not available for this question yet.
Question 11 MCQ · 2.0 marks
Suppose you perform the perceptron algorithm on the following dataset:
[[IMAGE:8f61ac644f74e0fb_8_63]]
Assume the initial weights [[IMAGE:8f61ac644f74e0fb_8_64]] and we pass the data points in the order given in the
table. The following rule is used for the classification:
[[IMAGE:8f61ac644f74e0fb_8_65]]
Based on the above data, answer the given subquestions.
After one epoch, your friend claims that the perceptron weights no longer need updating, i.e., the
model has converged. Specify whether the statement is true or false.



True
False
A published solution is not available for this question yet.
Question 12 NAT · 3.0 marks
Consider a neural network for a binary classification problem with one input and one output.
There is one hidden layer with two neurons and an output layer. Both the layers uses sigmoid as
an activation function. Cross entropy loss is used.
Following are the values of weights and biases:
[[IMAGE:8f61ac644f74e0fb_10_70]]
[[IMAGE:8f61ac644f74e0fb_10_71]]
Based on the above data, answer the given subquestions.
Compute the predicted output [[IMAGE:8f61ac644f74e0fb_10_72]] for the given input [[IMAGE:8f61ac644f74e0fb_10_73]] . Enter the answer correct to two decimal
places.




A published solution is not available for this question yet.
Question 13 NAT · 4.0 marks
Consider a neural network for a binary classification problem with one input and one output.
There is one hidden layer with two neurons and an output layer. Both the layers uses sigmoid as
an activation function. Cross entropy loss is used.
Following are the values of weights and biases:
[[IMAGE:8f61ac644f74e0fb_10_70]]
[[IMAGE:8f61ac644f74e0fb_10_71]]
Based on the above data, answer the given subquestions.
Compute
[[IMAGE:8f61ac644f74e0fb_11_74]]
.
Enter the answer correct to two decimal places.



A published solution is not available for this question yet.
Question 14 MCQ · 3.0 marks
Consider a linear regression model [[IMAGE:8f61ac644f74e0fb_11_75]] trained using momentum-based mini-batch gradient
descent with the loss function:
[[IMAGE:8f61ac644f74e0fb_11_76]]
where [[IMAGE:8f61ac644f74e0fb_11_77]] is the batch size. The algorithm uses a mini-batch size of [[IMAGE:8f61ac644f74e0fb_11_78]] , learning rate [[IMAGE:8f61ac644f74e0fb_11_79]] , and
the momentum coefficient [[IMAGE:8f61ac644f74e0fb_11_80]] . The weights and velocity are initialized as [[IMAGE:8f61ac644f74e0fb_11_81]] and
[[IMAGE:8f61ac644f74e0fb_11_82]] . The update rules are given by:
[[IMAGE:8f61ac644f74e0fb_11_83]]
[[IMAGE:8f61ac644f74e0fb_11_84]]
The first mini-batch consists of the following samples:
[[IMAGE:8f61ac644f74e0fb_11_85]]
Based on the above data, answer the given subquestions.
Compute the gradient [[IMAGE:8f61ac644f74e0fb_12_86]] for this mini-batch.












[[IMAGE:8f61ac644f74e0fb_12_87]]

[[IMAGE:8f61ac644f74e0fb_12_88]]

[[IMAGE:8f61ac644f74e0fb_12_89]]

[[IMAGE:8f61ac644f74e0fb_12_90]]

A published solution is not available for this question yet.
Question 15 NAT · 3.0 marks
Consider a linear regression model [[IMAGE:8f61ac644f74e0fb_11_75]] trained using momentum-based mini-batch gradient
descent with the loss function:
[[IMAGE:8f61ac644f74e0fb_11_76]]
where [[IMAGE:8f61ac644f74e0fb_11_77]] is the batch size. The algorithm uses a mini-batch size of [[IMAGE:8f61ac644f74e0fb_11_78]] , learning rate [[IMAGE:8f61ac644f74e0fb_11_79]] , and
the momentum coefficient [[IMAGE:8f61ac644f74e0fb_11_80]] . The weights and velocity are initialized as [[IMAGE:8f61ac644f74e0fb_11_81]] and
[[IMAGE:8f61ac644f74e0fb_11_82]] . The update rules are given by:
[[IMAGE:8f61ac644f74e0fb_11_83]]
[[IMAGE:8f61ac644f74e0fb_11_84]]
The first mini-batch consists of the following samples:
[[IMAGE:8f61ac644f74e0fb_11_85]]
Based on the above data, answer the given subquestions.
Using the gradient computed in the previous question, determine the velocity [[IMAGE:8f61ac644f74e0fb_12_91]] after the first
iteration.
Enter the answer correct to one decimal place.












A published solution is not available for this question yet.