cs3004_2024T1_Q1_NA.pdf
Deep Learning · Quiz 1 · Jan 2024
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 43 MSQ · 4.0 marks
Consider a neuron with binary inputs x1 and x2, and an output y. The neuron computes the
weighted sum of its inputs and produces an output according to a threshold. The threshold is
denoted as θ. The activation function is such that y = 1 if the weighted sum is greater than or equal
to θ, otherwise y = 0.
Which of the following statements are correct regarding the neuron’s ability to represent logical
AND and OR functions?
The neuron can implement the AND function by setting appropriate weights
and a threshold.
The neuron can implement the OR function by setting appropriate weights and
a threshold.
There exists a single set of weights and a threshold that allows the same
neuron to correctly implement both the AND and OR functions simultaneously.
Neurons are limited to implementing either the AND or the OR function and
cannot represent both simultaneously.
A published solution is not available for this question yet.
Question 44 NAT · 3.0 marks
Suppose we have a perceptron with two inputs, x1 and x2. This perceptron undergoes training on a
small dataset containing three points: (−1, 2) labeled as class 0, (0,−1) labeled as class 1, and (2, 1)
labeled as class 0. The weights of the perceptron are initialized to zeros, and the model is trained
until it reaches convergence. Given this scenario, what would be the assigned output class by the
trained perceptron for the new point (−2, 0)?
A published solution is not available for this question yet.
Question 45 NAT · 3.0 marks
You are training a neural network for sentiment analysis on a dataset of 10,000 text reviews. The
dataset is divided into 80% for training and 20% for testing. You decide to use Minibatch Gradient
Descent with a batch size of 32. If you perform a total of 100 epochs, how many parameter
updates will be performed in total?
A published solution is not available for this question yet.
Question 46 MCQ · 2.0 marks
Let’s assume a continuous function f(x1,x2) is approximated using 3D tower function with a 3
hidden layer neural network using 100 towers. How many minimum number of neurons will we
need to approximate the function f(x1,x2) ?
201
300
801
701
800
A published solution is not available for this question yet.
Question 47 NAT · 4.0 marks
[[IMAGE:d8a468ed6527ce80_6_1]]

A published solution is not available for this question yet.
Question 48 NAT · 4.0 marks
[[IMAGE:d8a468ed6527ce80_6_2]]

A published solution is not available for this question yet.
Question 49 MCQ · 4.0 marks
[[IMAGE:d8a468ed6527ce80_7_3]]

Increasing the value of b shifts the sigmoid function to the left (i.e., towards
negative infinity)
Increasing the value of b shifts the sigmoid function to the right (i.e., towards
positive infinity)
Decreasing the value of w increases the steepness of the sigmoid function
Increasing the value of w decreases the steepness of the sigmoid function
A published solution is not available for this question yet.
Question 50 MCQ · 4.0 marks
Which of the following is true, given the optimal learning rate?
Batch gradient descent is always guaranteed to converge to the global
optimum of a loss function.
Stochastic gradient descent is always guaranteed to converge to the global
optimum of a loss function.
For convex loss functions, stochastic gradient descent is guaranteed to
eventually converge to the global optimum while batch gradient descent is not.
For convex loss functions, both stochastic gradient descent and batch gradient
descent will eventually converge to the global optimum.
For convex loss functions, neither stochastic gradient descent nor batch
gradient descent are guaranteed to converge to the global optimum.
For convex loss functions, batch gradient descent is guaranteed to eventually
converge to the global optimum while stochastic gradient descent is not.
A published solution is not available for this question yet.
Question 51 NAT · 3.0 marks
[[IMAGE:d8a468ed6527ce80_8_4]]
Based on the above data, answer the given subquestions.
How many parameters (including biases) are there in the entire network?

A published solution is not available for this question yet.
Question 52 SHORT_TEXT · 4.0 marks
[[IMAGE:d8a468ed6527ce80_8_4]]
Based on the above data, answer the given subquestions.
Suppose that all elements in the input vector are zero and the corresponding true label is also 0.
Further, suppose that all the parameters are initialized to zero.
What is the loss value if cross-entropy loss is used? Use natural logarithm ln.

A published solution is not available for this question yet.
Question 53 SHORT_TEXT · 4.0 marks
[[IMAGE:d8a468ed6527ce80_8_4]]
Based on the above data, answer the given subquestions.
Assuming that all weights between layers **h3** and **O** are initialized to one, with no bias associated
with any neuron, what would be the computed cross-entropy loss for a given single data point? If
the provided information is insufficient, please enter −1.

A published solution is not available for this question yet.
Question 54 SHORT_TEXT · 4.0 marks
[[IMAGE:d8a468ed6527ce80_11_5]]
Based on the above data, answer the given subquestions.
Compute the cross entropy loss.
Note:If you think the given information is not sufficient to calculate the loss, then enter -1 as
answer.

A published solution is not available for this question yet.
Question 55 SHORT_TEXT · 4.0 marks
[[IMAGE:d8a468ed6527ce80_11_5]]
Based on the above data, answer the given subquestions.
[[IMAGE:d8a468ed6527ce80_12_6]]


A published solution is not available for this question yet.
Question 56 MCQ · 3.0 marks
Consider the following image:
[[IMAGE:d8a468ed6527ce80_13_7]]
As per your understanding of optimization algorithms, which of the following mappings will be
correct (assume optimal learning rate)?

1: Gradient Descent
2: Momentum based Gradient Descent
3: Nesterov Accelerated Gradient Descent
1: Momentum based Gradient Descent
2: Gradient Descent
3: Nesterov Accelerated Gradient Descent
1: Gradient Descent
2: Nesterov Accelerated Gradient Descent
3: Momentum based Gradient Descent
1: Momentum based Gradient Descent
2: Nesterov Accelerated Gradient Descent
3: Gradient Descent
**Programming in C**
**Section Id :** 64065351453
**Section Number :** 4
A published solution is not available for this question yet.