cs3004_2026T2_Q1_NA.pdf
Deep Learning · Quiz 1 · May 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 1.0 marks
State whether the following statement is true or false.
The decision boundary learned by a neural network is always non-linear.
True
False
Published solution
**1. Key concept:** A neural network's decision boundary is non-linear only if its activations are non-linear and there is at least one non-linear hidden layer.
**2. Counterexample:** A single neuron (perceptron or sigmoid), or a network whose layers all use linear/identity activations, gives a *linear* boundary. So 'always non-linear' is false.
**3. Conclude:** The statement is False.
Answer: B — False.
A: **Incorrect:** The boundary is not always non-linear. A single neuron or a purely linear network has a linear boundary.
B: **Correct:** A single neuron or a purely linear network has a linear boundary, so the statement is false.
Question 3 MCQ · 2.0 marks
Consider the following two statements and select the correct options.
**Statement I**: A single MP neuron can be used to represent all linearly separable Boolean
functions.
**Statement II**: Every linearly separable boolean function can be represented using at least one MP
neuron.
Only Statement I is true
Only Statement II is true
Both the statements are true.
None of the statements are true.
Published solution
**1. Statement I:** A single MP neuron has fixed unit weights (excitatory or inhibitory) and one threshold. It treats all excitatory inputs equally, so it cannot represent every linearly separable function. Example: \(f=x_1 \lor (x_2\land x_3)\) is linearly separable (\(2x_1+x_2+x_3\ge 2\)), but it is not symmetric in the inputs, so one MP neuron cannot compute it. Statement I is false.
**2. Statement II:** A network of MP neurons (one or more MP neurons) can represent any Boolean function, so it can represent every linearly separable Boolean function. Statement II is true.
**3. Conclude:** Only Statement II is true.
Answer: B — Only Statement II is true.
A: **Incorrect:** Statement I is false because one MP neuron cannot represent all linearly separable functions.
B: **Correct:** Statement II is true and Statement I is false.
C: **Incorrect:** Statement I is false, so both cannot be true.
D: **Incorrect:** Statement II is true, so 'none' is wrong.
Question 4 MCQ · 2.0 marks
Two runs of gradient descent on [[IMAGE:64ee263ead107bbe_3_2]] from the same starting point [[IMAGE:64ee263ead107bbe_3_3]] produce Graph
X and Graph Y using two learning rates:
[[IMAGE:64ee263ead107bbe_3_4]]
Which among the following options could be correct?



X uses [[IMAGE:64ee263ead107bbe_3_5]] , Y uses [[IMAGE:64ee263ead107bbe_3_6]]


X uses [[IMAGE:64ee263ead107bbe_3_7]] , Y uses [[IMAGE:64ee263ead107bbe_3_8]]


Published solution
**1. Given:** \(L(w)=w^2\), \(w_0=2\). The gradient is \(2w\), so the update is
\[w_{t+1}=w_t-\eta(2w_t)=(1-2\eta)w_t.\]
**2. Case \(\eta=0.3\):** The factor is \(1-0.6=0.4\), so \(|w|\) shrinks. Loss at step 1: \((0.4\times2)^2=0.64\). This matches Graph X, which drops from 4 to about 0.64 and then to 0.
**3. Case \(\eta=1.1\):** The factor is \(1-2.2=-1.2\), so \(|w|\) grows by 1.2 each step and the loss grows by \(1.44\) each step. After 10 steps, loss \(=4\times1.44^{10}\approx153\). This matches Graph Y.
**4. Conclude:** X uses \(\eta=0.3\) and Y uses \(\eta=1.1\).
Answer: B — X uses \(\eta=0.3\), Y uses \(\eta=1.1\).
A: **Incorrect:** The large rate 1.1 would make the loss diverge, so it cannot be Graph X.
B: **Correct:** Rate 0.3 converges (factor 0.4) and rate 1.1 diverges (factor -1.2).
Question 5 MCQ · 2.0 marks
The loss function [[IMAGE:64ee263ead107bbe_3_9]] is minimized starting at [[IMAGE:64ee263ead107bbe_3_10]] with [[IMAGE:64ee263ead107bbe_3_11]] using four
methods: Vanilla Gradient Descent ( [[IMAGE:64ee263ead107bbe_3_12]] ), and momentum based gradient descent with
[[IMAGE:64ee263ead107bbe_3_13]] .
[[IMAGE:64ee263ead107bbe_4_14]]
Match each curve to its [[IMAGE:64ee263ead107bbe_4_15]] value:







S: [[IMAGE:64ee263ead107bbe_4_16]] , T: [[IMAGE:64ee263ead107bbe_4_17]] , U: [[IMAGE:64ee263ead107bbe_4_18]] , V: [[IMAGE:64ee263ead107bbe_4_19]]




S: [[IMAGE:64ee263ead107bbe_4_20]] , T: [[IMAGE:64ee263ead107bbe_4_21]] , U: [[IMAGE:64ee263ead107bbe_4_22]] , V: [[IMAGE:64ee263ead107bbe_4_23]]




S: [[IMAGE:64ee263ead107bbe_4_24]] , T: [[IMAGE:64ee263ead107bbe_4_25]] , U: [[IMAGE:64ee263ead107bbe_4_26]] , V: [[IMAGE:64ee263ead107bbe_4_27]]




S: [[IMAGE:64ee263ead107bbe_4_28]] , T: [[IMAGE:64ee263ead107bbe_4_29]] , U: [[IMAGE:64ee263ead107bbe_4_30]] , V: [[IMAGE:64ee263ead107bbe_4_31]]




Published solution
**1. Given:** \(L=(w-5)^2\), \(\eta=0.04\). The curvature is \(\lambda=2\), so \(\eta\lambda=0.08\). The error \(e=w-5\) follows
\[e_{t+1}=(1+\beta-0.08)\,e_t-\beta\,e_{t-1}.\]
**2. Find the decay of each \(\beta\):** The characteristic equation is \(r^2-(0.92+\beta)r+\beta=0\).
- \(\beta=0\): \(r=0.92\). Smooth, monotone decay (slow).
- \(\beta=0.5\): real roots \(\approx0.77,\,0.65\). Faster monotone decay, no oscillation.
- \(\beta=0.9\): complex roots with \(|r|=\sqrt{0.9}\approx0.95\). Oscillating, decaying.
- \(\beta=0.99\): complex roots with \(|r|\approx0.995\). Large oscillations, decaying very slowly.
**3. Match the curves:** S (smooth, slow) is \(\beta=0\). T (fastest, no oscillation) is \(\beta=0.5\). U (oscillates, then dies out) is \(\beta=0.9\). V (big long-lasting oscillation) is \(\beta=0.99\).
Answer: A — S: \(\beta=0\), T: \(\beta=0.5\), U: \(\beta=0.9\), V: \(\beta=0.99\).
A: **Correct:** S is the smooth vanilla curve, T is fastest, U oscillates and settles, V oscillates the longest.
B: **Incorrect:** This reverses the order. The most oscillating curve V needs the largest beta, not 0.
C: **Incorrect:** T (no oscillation, fastest) cannot be vanilla GD, and V cannot be 0.9 since it oscillates the longest.
D: **Incorrect:** T has no oscillation, so it cannot be 0.9, and U oscillates more than 0.5 would.
Question 6 MCQ · 3.0 marks
A perceptron is being trained on a 2D dataset using the **Perceptron Learning Algorithm**. The current weight vector is \(w=\begin{bmatrix}3\\4\end{bmatrix}\).
For a misclassified training example \(x=\begin{bmatrix}4\\3\end{bmatrix}\), with true label \(y=-1\), the perceptron update rule with learning rate \(\eta=1\) is applied: \(w_{\text{new}}=w-x\).
Let \(\theta_{\text{new}}\) denote the angle between \(w_{\text{new}}\) and \(x\). Which of the following is equal to \(\theta_{\text{new}}\)?
[[IMAGE:64ee263ead107bbe_5_41]]

[[IMAGE:64ee263ead107bbe_5_42]]

[[IMAGE:64ee263ead107bbe_5_43]]

[[IMAGE:64ee263ead107bbe_5_44]]

Published solution
**1. Given:** \(w=\begin{bmatrix}3\\4\end{bmatrix}\), \(x=\begin{bmatrix}4\\3\end{bmatrix}\), \(w_{new}=w-x\).
**2. Compute \(w_{new}\):**
\[w_{new}=\begin{bmatrix}3-4\\4-3\end{bmatrix}=\begin{bmatrix}-1\\1\end{bmatrix}\]
**3. Use the angle formula:** \(\cos\theta=\dfrac{w_{new}\cdot x}{\|w_{new}\|\,\|x\|}\).
\[w_{new}\cdot x=-4+3=-1,\quad \|w_{new}\|=\sqrt2,\quad \|x\|=5\]
\[\cos\theta_{new}=\frac{-1}{5\sqrt2}\]
**4. Conclude:** \(\theta_{new}=\cos^{-1}\!\left(\dfrac{-1}{5\sqrt2}\right)\).
Answer: C — \(\cos^{-1}\!\left(\dfrac{-1}{5\sqrt2}\right)\).
A: **Incorrect:** The dot product is -1, not 24.
B: **Incorrect:** The dot product is -1, so the cosine is negative, not positive.
C: **Correct:** The dot product is -1 and the norms are \(\sqrt2\) and 5, giving \(-1/(5\sqrt2)\).
D: **Incorrect:** The vectors are not parallel, so the angle is not 0°.
Question 7 MSQ · 3.0 marks
Which of the following statements are TRUE about training neural networks?
Stochastic gradient descent (batch size [[IMAGE:64ee263ead107bbe_5_45]] ) computes the true gradient of
the total loss function at each step.

In mini-batch gradient descent with batch size [[IMAGE:64ee263ead107bbe_5_46]] on a dataset of [[IMAGE:64ee263ead107bbe_5_47]] points,
one epoch consists of [[IMAGE:64ee263ead107bbe_5_48]] parameter updates.



Momentum-based gradient descent can overshoot a narrow valley because
accumulated velocity carries the parameters past the minimum.
In vanilla (batch) gradient descent, each step is guaranteed to decrease the
loss for a sufficiently small learning rate on a smooth loss function.
Published solution
**1. A:** SGD with \(B=1\) uses the gradient of a *single* example. This is a noisy estimate, not the true gradient of the total loss. False.
**2. B:** One epoch means one pass over all \(N\) points. With batch size \(B\) there are \(\lceil N/B\rceil\) batches, so \(\lceil N/B\rceil\) updates. True.
**3. C:** Momentum keeps accumulated velocity, so it can run past the minimum in a narrow valley and oscillate. True.
**4. D:** For a smooth loss and small enough \(\eta\), a full-batch gradient step decreases the loss (descent lemma). True.
**5. Conclude:** B, C and D are true.
Answer: B, C, D — B, C, D are true.
A: **Incorrect:** SGD uses one example, so it gives only a noisy estimate of the true gradient.
B: **Correct:** One epoch has \(\lceil N/B\rceil\) batches, so that many updates.
C: **Correct:** Accumulated velocity can carry the parameters past the minimum.
D: **Correct:** With a small enough learning rate on a smooth loss, each full-batch step lowers the loss.
Question 8 MSQ · 3.0 marks
Consider a hidden layer with [[IMAGE:64ee263ead107bbe_6_49]] neurons, all initialized with identical weight vectors [[IMAGE:64ee263ead107bbe_6_50]] and identical
bias [[IMAGE:64ee263ead107bbe_6_51]] . The activation function is sigmoid. Which of the following are TRUE?



After one gradient descent step, the neurons will have different weights.
For any input [[IMAGE:64ee263ead107bbe_6_52]] , all [[IMAGE:64ee263ead107bbe_6_53]] neurons will produce the same output.


During backpropagation, the gradients of all [[IMAGE:64ee263ead107bbe_6_54]] neurons are always zero
because their weights are identical.

This layer is functionally equivalent to having a single neuron, because all [[IMAGE:64ee263ead107bbe_6_55]]
neurons compute the same function.

Published solution
**1. Key idea:** All neurons have identical weights and bias, so they compute the same output and receive the same gradient. This is the *symmetry problem*.
**2. A:** The gradients are identical, so after one step the weights stay identical. False.
**3. B:** Same \(w\), same \(b\), same input \(\Rightarrow\) same output \(\sigma(w^Tx+b)\). True.
**4. C:** The gradients are identical across neurons, but not zero in general. False.
**5. D:** All \(n\) neurons compute the same function, so the layer acts like one neuron (with combined outgoing weight). True.
**6. Conclude:** B and D are true.
Answer: B, D — B, D.
A: **Incorrect:** The gradients are identical, so the weights remain identical after the step.
B: **Correct:** Identical weights, bias and input give identical outputs.
C: **Incorrect:** The gradients are identical but generally non-zero.
D: **Correct:** All neurons compute the same function, so one neuron is equivalent.
Question 9 MSQ · 3.0 marks
Consider an MP neuron with three inputs
[[IMAGE:64ee263ead107bbe_6_56]]
Inputs [[IMAGE:64ee263ead107bbe_6_57]] and [[IMAGE:64ee263ead107bbe_6_58]] are excitatory, while [[IMAGE:64ee263ead107bbe_6_59]] is inhibitory. The threshold of the neuron is 2.The neuron
fires (outputs 1) if:
● the inhibitory input is inactive, and
● the sum of excitatory inputs is at least the threshold.
Which of the following statements are correct?




For input [[IMAGE:64ee263ead107bbe_6_60]] , the neuron outputs 1.

For input [[IMAGE:64ee263ead107bbe_6_61]] , the neuron outputs 1.

For input [[IMAGE:64ee263ead107bbe_7_62]] , the neuron outputs 0.

Increasing the threshold from 2 to 3 would not change the output of the
neuron for any input combination.
Published solution
**1. Rule:** Output \(=1\) only if \(x_3=0\) and \(x_1+x_2\ge 2\).
**2. Check inputs:**
- \((1,1,0)\): \(x_3=0\), sum \(=2\ge2\), so output 1. A is true.
- \((1,1,1)\): \(x_3=1\) (inhibitor active), so output 0. B is false.
- \((1,0,0)\): sum \(=1<2\), so output 0. C is true.
- Threshold 3: max excitatory sum is 2, so the neuron never fires. Input \((1,1,0)\) changes from 1 to 0. D is false.
**3. Conclude:** A and C are correct.
Answer: A, C — A, C.
A: **Correct:** Inhibitory input is 0 and the sum of excitatory inputs is 2, so it fires.
B: **Incorrect:** The inhibitory input is active, so the output is 0.
C: **Correct:** The sum of excitatory inputs is 1, which is below 2, so the output is 0.
D: **Incorrect:** With threshold 3 the neuron never fires, so the output of (1,1,0) changes from 1 to 0.
Question 10 MSQ · 4.0 marks
Consider a perceptron with initial parameters
[[IMAGE:64ee263ead107bbe_7_63]]
The perceptron predicts class [[IMAGE:64ee263ead107bbe_7_64]] if
[[IMAGE:64ee263ead107bbe_7_65]]
and class [[IMAGE:64ee263ead107bbe_7_66]] otherwise. Whenever a training example [[IMAGE:64ee263ead107bbe_7_67]] is misclassified, where
[[IMAGE:64ee263ead107bbe_7_68]] , the perceptron updates its parameters as
[[IMAGE:64ee263ead107bbe_7_69]]
The following training examples are processed once in the order
[[IMAGE:64ee263ead107bbe_7_70]]
[[IMAGE:64ee263ead107bbe_8_71]]
Which of the following statements are correct?









The first weight update occurs after processing point [[IMAGE:64ee263ead107bbe_8_72]] .

The second weight update occurs after processing point [[IMAGE:64ee263ead107bbe_8_73]] .

The final bias equals [[IMAGE:64ee263ead107bbe_8_74]] .

The perceptron performs exactly two weight updates.
Published solution
**1. Start:** \(w=(2,-1)\), \(b=0\). Predict \(+1\) if \(w^Tx+b\ge0\).
**2. Point A \((3,1),+1\):** \(6-1+0=5\ge0\Rightarrow+1\). Correct, no update.
**3. Point B \((2,0),-1\):** \(4\ge0\Rightarrow+1\). Wrong, update (first update):
\[w=(2,-1)-(2,0)=(0,-1),\quad b=0-1=-1\]
**4. Point C \((1,2),-1\):** \(0\cdot1-1\cdot2-1=-3<0\Rightarrow-1\). Correct, no update.
**5. Point D \((4,4),+1\):** \(0-4-1=-5<0\Rightarrow-1\). Wrong, update (second update):
\[w=(0,-1)+(4,4)=(4,3),\quad b=-1+1=0\]
**6. Check options:** First update after B (true). Second update after D, not C (false). Final bias is 0, not -1 (false). Exactly two updates (true).
**7. Conclude:** A and D are correct.
Answer: A, D — A, D.
A: **Correct:** The first mistake is at B.
B: **Incorrect:** C is classified correctly; the second update happens at D.
C: **Incorrect:** The final bias is -1+1 = 0.
D: **Correct:** Updates happen only at B and D.
Question 11 MSQ · 4.0 marks
Which of the following statements regarding perceptrons and sigmoid neurons are correct?
The output of a perceptron changes abruptly when the weighted sum of
inputs crosses the threshold, whereas the output of a sigmoid neuron changes smoothly.
The sigmoid activation function is differentiable, making it suitable for
gradient-based learning algorithms.
By choosing sufficiently large weights, the output of a sigmoid neuron can
closely approximate the output of a perceptron.
A multilayer network of perceptrons with a single hidden layer can represent
any Boolean function exactly.
A multilayer network of sigmoid neurons with a single hidden layer can
approximate any continuous function to any desired precision, provided enough hidden neurons
are available.
A single sigmoid neuron can represent every non-linear decision boundary in
[[IMAGE:64ee263ead107bbe_8_75]] .

Published solution
**1. A:** A perceptron is a hard step function, so its output jumps at the threshold. A sigmoid changes smoothly. True.
**2. B:** The sigmoid is differentiable everywhere, so gradient-based learning works. True.
**3. C:** For \(\sigma(kz)\) with large \(k\), the curve approaches a step function. True.
**4. D:** One hidden layer of perceptrons (one for each input pattern) can represent any Boolean function. True.
**5. E:** This is the universal approximation theorem. True.
**6. F:** A single sigmoid neuron has a *linear* decision boundary \(w^Tx+b=0\). It cannot represent every non-linear boundary. False.
**7. Conclude:** A, B, C, D and E are correct.
Answer: A, B, C, D, E — A, B, C, D, E.
A: **Correct:** Perceptron output jumps; sigmoid output varies smoothly.
B: **Correct:** The sigmoid has a derivative everywhere, which allows gradient-based learning.
C: **Correct:** Large weights make the sigmoid very steep, close to a step.
D: **Correct:** A hidden layer of perceptrons can realise any Boolean function.
E: **Correct:** This is the universal approximation theorem.
F: **Incorrect:** A single sigmoid neuron has a linear boundary, so it cannot represent all non-linear boundaries.
Question 12 MSQ · 4.0 marks
Which of the following statements are true regarding the sigmoid function and its Taylor series
approximation?
Let
[[IMAGE:64ee263ead107bbe_9_76]]
be the sigmoid activation function.

Using the first-order Taylor approximation [[IMAGE:64ee263ead107bbe_9_77]] , the sigmoid function
near [[IMAGE:64ee263ead107bbe_9_78]] can be approximated as [[IMAGE:64ee263ead107bbe_9_79]] .



The derivative of the sigmoid function at [[IMAGE:64ee263ead107bbe_9_80]] is [[IMAGE:64ee263ead107bbe_9_81]] .


The first-order Taylor approximation implies that the sigmoid function is
exactly linear for all values of [[IMAGE:64ee263ead107bbe_9_82]] .

As [[IMAGE:64ee263ead107bbe_9_83]] becomes very large, the linear approximation obtained from the Taylor
series becomes increasingly inaccurate.

The smooth differentiability of the sigmoid function enables the use of
gradient-based learning algorithms, unlike the hard-threshold perceptron.
Published solution
**1. A:** \(\sigma(z)\approx\dfrac{1}{1+(1-z)}=\dfrac{1}{2-z}=\dfrac12\cdot\dfrac1{1-z/2}\approx\dfrac12+\dfrac z4\) for small \(z\). True (this equals the Taylor series of \(\sigma\) at 0).
**2. B:** \(\sigma'(z)=\sigma(z)(1-\sigma(z))\), so \(\sigma'(0)=\tfrac12\cdot\tfrac12=\tfrac14\). True.
**3. C:** A first-order approximation is valid only near \(z=0\). \(\sigma\) is bounded in \((0,1)\), but the line \(\tfrac12+\tfrac z4\) is unbounded. False.
**4. D:** For large \(|z|\), the line goes outside \([0,1]\) while \(\sigma\) saturates, so the error grows. True.
**5. E:** The sigmoid is smooth and differentiable, unlike the hard-threshold perceptron. True.
**6. Conclude:** A, B, D and E are true.
Answer: A, B, D, E — A, B, D, E.
A: **Correct:** Expanding near 0 gives 1/2 + z/4.
B: **Correct:** \(\sigma'(0)=\sigma(0)(1-\sigma(0))=1/4\).
C: **Incorrect:** The sigmoid is not linear; the approximation is valid only near z = 0.
D: **Correct:** The line is unbounded while the sigmoid saturates, so the error grows with |z|.
E: **Correct:** The sigmoid is differentiable, so gradient-based learning works, unlike the step function.
Question 13 NAT · 3.0 marks
Consider a regression problem of predicting the price of a house using two features: Overall
quality [[IMAGE:64ee263ead107bbe_9_84]] , living area (sq. ft) [[IMAGE:64ee263ead107bbe_9_85]] . A simple feed forward neural network is used. It has three
layers:
● The input layer consisting of the two features [[IMAGE:64ee263ead107bbe_10_86]] .
● A hidden layer consisting of three sigmoid (logistic) neurons with weights [[IMAGE:64ee263ead107bbe_10_87]] and
biases [[IMAGE:64ee263ead107bbe_10_88]] .
● Output layer with a single linear neuron that outputs the predicted house value [[IMAGE:64ee263ead107bbe_10_89]] with weights
[[IMAGE:64ee263ead107bbe_10_90]] and bias [[IMAGE:64ee263ead107bbe_10_91]] .
[[IMAGE:64ee263ead107bbe_10_92]]
Suppose we start the Gradient Descent algorithm by setting all the weights [[IMAGE:64ee263ead107bbe_10_93]] equal to 0, [[IMAGE:64ee263ead107bbe_10_94]]
equal to 100, and the biases equal to 0. What does the initial network give for the price of a house
with overall quality equal to 8 and living area equal to 3000 square feet?











Published solution
**1. Given:** \(W_1=0\), \(b_1=0\), \(W_2=[100,100,100]\), \(b_2=0\), input \(x=(8,3000)\).
**2. Hidden layer:** \(a_1=W_1x+b_1=0\), so each hidden output is \(h_1=\sigma(0)=0.5\).
**3. Output layer (linear):**
\[\hat y=W_2h_1+b_2=100(0.5)+100(0.5)+100(0.5)=150\]
**4. Conclude:** The initial prediction does not depend on the input values here.
Answer: 150.
Question 14 NAT · 3.0 marks
Consider a 3-class softmax output with pre-activations [[IMAGE:64ee263ead107bbe_10_95]] .
Find the value of [[IMAGE:64ee263ead107bbe_10_96]] for which the predicted probability of class 1 is exactly [[IMAGE:64ee263ead107bbe_10_97]] .
Enter the answer correct to two decimal places.



Published solution
**1. Given:** \(a_L=[a_1,0,0]\) and \(\hat y_1=0.5\).
**2. Softmax:**
\[\hat y_1=\frac{e^{a_1}}{e^{a_1}+e^0+e^0}=\frac{e^{a_1}}{e^{a_1}+2}\]
**3. Solve:**
\[\frac{e^{a_1}}{e^{a_1}+2}=0.5\Rightarrow e^{a_1}=0.5e^{a_1}+1\Rightarrow e^{a_1}=2\Rightarrow a_1=\ln2\approx0.69\]
Answer: \(a_1=\ln 2\approx0.69\).
Question 15 NAT · 3.0 marks
Consider a sigmoid neuron
[[IMAGE:64ee263ead107bbe_11_98]]
For a training example [[IMAGE:64ee263ead107bbe_11_99]] , let
[[IMAGE:64ee263ead107bbe_11_100]]
The squared-error loss is given by
[[IMAGE:64ee263ead107bbe_11_101]]
Find the value of [[IMAGE:64ee263ead107bbe_11_102]] . Round your answer to two decimal places.





Published solution
**1. Given:** \(x=2\), \(y=0.8\), \(w=b=0\), \(L=\tfrac12(f-y)^2\).
**2. Forward pass:** \(f=\sigma(0\cdot2+0)=\sigma(0)=0.5\).
**3. Chain rule:**
\[\nabla_wL=(f-y)\,f(1-f)\,x\]
**4. Substitute:**
\[\nabla_wL=(0.5-0.8)(0.5)(0.5)(2)=(-0.3)(0.25)(2)=-0.15\]
Answer: \(-0.15\).
Question 16 MCQ · 3.0 marks
You build a feedforward neural network to predict which of 10 categories a product belongs to,
based on 50 numerical features (price, weight, customer rating, etc.). The network has the
following architecture:
Input [[IMAGE:64ee263ead107bbe_12_103]] [[IMAGE:64ee263ead107bbe_12_104]] Hidden Layer 1 ( [[IMAGE:64ee263ead107bbe_12_105]] sigmoid neurons) [[IMAGE:64ee263ead107bbe_12_106]] Hidden Layer 2 ( [[IMAGE:64ee263ead107bbe_12_107]] sigmoid neurons)
[[IMAGE:64ee263ead107bbe_12_108]] Output ( [[IMAGE:64ee263ead107bbe_12_109]] softmax neurons)
with cross-entropy loss [[IMAGE:64ee263ead107bbe_12_110]] , where [[IMAGE:64ee263ead107bbe_12_111]] is the true class.
[[IMAGE:64ee263ead107bbe_12_112]]
[[IMAGE:64ee263ead107bbe_12_113]]
[[IMAGE:64ee263ead107bbe_12_114]]
[[IMAGE:64ee263ead107bbe_12_115]]
[[IMAGE:64ee263ead107bbe_12_116]]
Based on the above data, answer the given subquestions.
Suppose you use the identity function ( [[IMAGE:64ee263ead107bbe_12_117]] ) in both the hidden layers, while keeping softmax
at the output. Which of the following is TRUE?















The network can learn non-linear decision boundaries because the softmax
output is non-linear.
The network is equivalent to a single softmax regression [[IMAGE:64ee263ead107bbe_12_118]]
where [[IMAGE:64ee263ead107bbe_12_119]] .


The network will fail to train because gradients cannot flow through linear
layers.
Training will force [[IMAGE:64ee263ead107bbe_13_120]] , [[IMAGE:64ee263ead107bbe_13_121]] , and [[IMAGE:64ee263ead107bbe_13_122]] to converge to identity matrices, making
the hidden layers redundant.



Published solution
**1. Setup:** With \(h_i=a_i\) in both hidden layers,
\[a_3=W_3(W_2(W_1x+b_1)+b_2)+b_3=W_3W_2W_1x+(W_3W_2b_1+W_3b_2+b_3).\]
**2. Meaning:** So \(a_3=W'x+b'\) with \(W'=W_3W_2W_1\). The hidden layers add no non-linearity, and the model is one softmax regression.
**3. Check the others:** Softmax only normalises the scores; the boundaries remain linear. Gradients flow normally through linear layers. Nothing forces the weights to become identity matrices.
Answer: B — The network is equivalent to a single softmax regression with \(W'=W_3W_2W_1\).
A: **Incorrect:** The decision boundary is still linear, since softmax does not add curvature to the boundaries between classes.
B: **Correct:** Stacked linear layers collapse into one linear map \(W'=W_3W_2W_1\), followed by softmax.
C: **Incorrect:** Gradients flow fine through linear layers, so training works.
D: **Incorrect:** Nothing pushes the weights toward identity matrices.
Question 17 NAT · 3.0 marks
You build a feedforward neural network to predict which of 10 categories a product belongs to,
based on 50 numerical features (price, weight, customer rating, etc.). The network has the
following architecture:
Input [[IMAGE:64ee263ead107bbe_12_103]] [[IMAGE:64ee263ead107bbe_12_104]] Hidden Layer 1 ( [[IMAGE:64ee263ead107bbe_12_105]] sigmoid neurons) [[IMAGE:64ee263ead107bbe_12_106]] Hidden Layer 2 ( [[IMAGE:64ee263ead107bbe_12_107]] sigmoid neurons)
[[IMAGE:64ee263ead107bbe_12_108]] Output ( [[IMAGE:64ee263ead107bbe_12_109]] softmax neurons)
with cross-entropy loss [[IMAGE:64ee263ead107bbe_12_110]] , where [[IMAGE:64ee263ead107bbe_12_111]] is the true class.
[[IMAGE:64ee263ead107bbe_12_112]]
[[IMAGE:64ee263ead107bbe_12_113]]
[[IMAGE:64ee263ead107bbe_12_114]]
[[IMAGE:64ee263ead107bbe_12_115]]
[[IMAGE:64ee263ead107bbe_12_116]]
Based on the above data, answer the given subquestions.
On a training example with true class [[IMAGE:64ee263ead107bbe_13_123]] , the network's softmax output is [[IMAGE:64ee263ead107bbe_13_124]] . What is the
[[IMAGE:64ee263ead107bbe_13_125]]
value of ? Enter the answer correct to two decimal places.

















Published solution
**1. Concept:** For softmax with cross-entropy loss \(\mathcal L=-\log\hat y_\ell\), the gradient with respect to the pre-activations is
\[\frac{\partial\mathcal L}{\partial a_{L,i}}=\hat y_i-\mathbb 1[i=\ell].\]
**2. Substitute for \(i=\ell\):**
\[\frac{\partial\mathcal L}{\partial a_{L,\ell}}=\hat y_\ell-1=0.99-1=-0.01\]
Answer: \(-0.01\).
Question 18 NAT · 2.0 marks
You are training a neural network to classify 200 species of birds from audio recordings. You have
collected a dataset of [[IMAGE:64ee263ead107bbe_13_126]] audio clips. You decide to start with vanilla gradient descent.
Based on the above data, answer the given subquestions.
With vanilla gradient descent, how many parameter updates does the network perform per
epoch?

Published solution
**1. Concept:** Vanilla (full-batch) gradient descent computes the gradient over the whole dataset before each update.
**2. Count:** One pass over all 100,000 clips (one epoch) gives exactly one update.
Answer: 1.
Question 19 NAT · 2.0 marks
You are training a neural network to classify 200 species of birds from audio recordings. You have
collected a dataset of [[IMAGE:64ee263ead107bbe_13_126]] audio clips. You decide to start with vanilla gradient descent.
Based on the above data, answer the given subquestions.
You switch to mini-batch gradient descent with batch size [[IMAGE:64ee263ead107bbe_14_127]] . How many parameter updates
does the network now perform per epoch?


Published solution
**1. Given:** \(N=100{,}000\) clips, batch size \(B=50\).
**2. Formula:** Updates per epoch \(=\lceil N/B\rceil\).
**3. Substitute:**
\[\frac{100000}{50}=2000\]
Answer: 2000.