cs3004_2026T1_ET_FN.pdf
Deep Learning · End Term · Jan 2026 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 2.0 marks
Consider a two-dimensional image of size [[IMAGE:0bc72138559e0d4e_2_2]] and a square filter of size [[IMAGE:0bc72138559e0d4e_2_3]] . How much
padding would you need on each side to keep the output height and width equal to the input
height and width? Assume a stride of 1.


[[IMAGE:0bc72138559e0d4e_2_4]]

[[IMAGE:0bc72138559e0d4e_2_5]]

[[IMAGE:0bc72138559e0d4e_2_6]]

[[IMAGE:0bc72138559e0d4e_2_7]]

A published solution is not available for this question yet.
Question 3 MCQ · 2.0 marks
Consider a vanilla Recurrent Neural Network (RNN) and an LSTM network with the same input
dimension, hidden dimension, and output dimension.
**Statement 1:** The number of trainable parameters in a vanilla RNN is higher than that in an LSTM
network.
**Statement 2:** An LSTM has a higher number of parameters because it contains multiple gates
(input gate, forget gate, output gate, and candidate state), each having separate weight matrices
and biases.
Choose the correct option from the following.
Statement 1 is true, and Statement 2 is the correct reason.
Statement 1 is true, but Statement 2 is false.
Statement 1 is false, but Statement 2 is true.
Statement 1 is false, and Statement 2 is false.
A published solution is not available for this question yet.
Question 4 MCQ · 3.0 marks
Which of the functions given below has the steepest slope at the point [[IMAGE:0bc72138559e0d4e_3_8]] ?

[[IMAGE:0bc72138559e0d4e_3_9]]

[[IMAGE:0bc72138559e0d4e_3_10]]

[[IMAGE:0bc72138559e0d4e_3_11]]

[[IMAGE:0bc72138559e0d4e_3_12]]

A published solution is not available for this question yet.
Question 5 MCQ · 3.0 marks
Consider a XOR function for input [[IMAGE:0bc72138559e0d4e_3_13]] , where the points [[IMAGE:0bc72138559e0d4e_3_14]] belong to one class
and [[IMAGE:0bc72138559e0d4e_3_15]] belong to another class. It is not possible to separate these two classes by
using a linear decision boundary. We consider a neural network with:
• Input [[IMAGE:0bc72138559e0d4e_3_16]]
• One hidden layer with 2 neurons and ReLU activation
• Output layer: one neuron with linear activation
[[IMAGE:0bc72138559e0d4e_4_17]]
The network is defined as:
[[IMAGE:0bc72138559e0d4e_4_18]]
Consider the weight matrices:
[[IMAGE:0bc72138559e0d4e_4_19]]
Can this network correctly classify XOR using a threshold at [[IMAGE:0bc72138559e0d4e_4_20]] ?








Yes
No
A published solution is not available for this question yet.
Question 6 MCQ · 3.0 marks
Match the attention mechanism in Column I with the correct description in Column II.
[[IMAGE:0bc72138559e0d4e_4_21]]

[[IMAGE:0bc72138559e0d4e_4_22]]

[[IMAGE:0bc72138559e0d4e_5_23]]

[[IMAGE:0bc72138559e0d4e_5_24]]

[[IMAGE:0bc72138559e0d4e_5_25]]

A published solution is not available for this question yet.
Question 7 NAT · 2.0 marks
Consider a neural network with the following architecture:
• Input: [[IMAGE:0bc72138559e0d4e_5_26]] , Output: [[IMAGE:0bc72138559e0d4e_5_27]] .
• [[IMAGE:0bc72138559e0d4e_5_28]] hidden layers with sigmoid(logistic) as the activation function, where each hidden layer has [[IMAGE:0bc72138559e0d4e_5_29]]
neurons. One neuron in the output layer with sigmoid as the activation function.
• All the weights are initialized as zero. Ignore the biases.
Based on the above data, answer the given subquestions.
What will be the output [[IMAGE:0bc72138559e0d4e_5_30]] ?
Enter the answer correct to one decimal place.





A published solution is not available for this question yet.
Question 8 NAT · 2.0 marks
Consider a neural network with the following architecture:
• Input: [[IMAGE:0bc72138559e0d4e_5_26]] , Output: [[IMAGE:0bc72138559e0d4e_5_27]] .
• [[IMAGE:0bc72138559e0d4e_5_28]] hidden layers with sigmoid(logistic) as the activation function, where each hidden layer has [[IMAGE:0bc72138559e0d4e_5_29]]
neurons. One neuron in the output layer with sigmoid as the activation function.
• All the weights are initialized as zero. Ignore the biases.
Based on the above data, answer the given subquestions.
What will be the output [[IMAGE:0bc72138559e0d4e_6_31]] if we change the activation function in the outer layer to linear?





A published solution is not available for this question yet.
Question 9 MSQ · 3.0 marks
Consider a neural network with the following architecture:
• Input: [[IMAGE:0bc72138559e0d4e_5_26]] , Output: [[IMAGE:0bc72138559e0d4e_5_27]] .
• [[IMAGE:0bc72138559e0d4e_5_28]] hidden layers with sigmoid(logistic) as the activation function, where each hidden layer has [[IMAGE:0bc72138559e0d4e_5_29]]
neurons. One neuron in the output layer with sigmoid as the activation function.
• All the weights are initialized as zero. Ignore the biases.
Based on the above data, answer the given subquestions.
Suppose you perform the backpropagation using the original configuration of the network and
update the parameters of the network using stochastic gradient descent. Let [[IMAGE:0bc72138559e0d4e_6_32]] be the weight
associated with the first hidden layer to input. Then, which among the following will be true for
[[IMAGE:0bc72138559e0d4e_6_33]] ?






It remains the same after the first update using stochastic gradient descent.
Different neurons in the first hidden layer will receive different gradient
updates.
All the elements in [[IMAGE:0bc72138559e0d4e_6_34]] will be zero after the first update.

Entries in [[IMAGE:0bc72138559e0d4e_6_35]] is either all positive or all negative.

A published solution is not available for this question yet.
Question 10 NAT · 2.0 marks
For a binary classification task, consider an image [[IMAGE:0bc72138559e0d4e_6_36]] of size [[IMAGE:0bc72138559e0d4e_6_37]] passed through a toy
convolutional network that has the following architecture:
• **ConvLayer:** filter size = [[IMAGE:0bc72138559e0d4e_7_38]] , and number of filter/channels = 1, with stride 1 and no padding.
• Convolved output is passed through a ReLU activation function.
• **Max-Pool Layer:** filter size = [[IMAGE:0bc72138559e0d4e_7_39]] , with stride 2 and padding not applied.
• Output of Max-pool layer is taken as [[IMAGE:0bc72138559e0d4e_7_40]] .
• Ignore the bias.
Following are the image and the kernel matrix [[IMAGE:0bc72138559e0d4e_7_41]] :
[[IMAGE:0bc72138559e0d4e_7_42]]
Based on the above data, answer the given subquestions.
Find the total number of parameters in the network.







A published solution is not available for this question yet.
Question 11 NAT · 3.0 marks
For a binary classification task, consider an image [[IMAGE:0bc72138559e0d4e_6_36]] of size [[IMAGE:0bc72138559e0d4e_6_37]] passed through a toy
convolutional network that has the following architecture:
• **ConvLayer:** filter size = [[IMAGE:0bc72138559e0d4e_7_38]] , and number of filter/channels = 1, with stride 1 and no padding.
• Convolved output is passed through a ReLU activation function.
• **Max-Pool Layer:** filter size = [[IMAGE:0bc72138559e0d4e_7_39]] , with stride 2 and padding not applied.
• Output of Max-pool layer is taken as [[IMAGE:0bc72138559e0d4e_7_40]] .
• Ignore the bias.
Following are the image and the kernel matrix [[IMAGE:0bc72138559e0d4e_7_41]] :
[[IMAGE:0bc72138559e0d4e_7_42]]
Based on the above data, answer the given subquestions.
[[IMAGE:0bc72138559e0d4e_7_44]]
[[IMAGE:0bc72138559e0d4e_7_45]]
The loss is defined as [[IMAGE:0bc72138559e0d4e_7_43]] . If then compute .










A published solution is not available for this question yet.
Question 12 MSQ · 2.0 marks
After training a neural network, the training error is observed to be 10%. It gives the test error to
be 56%. Which of the following methods can be used to reduce the test error?
Adam
ReLu activation
Injecting noise at input
Maxout
L2 regularization
A published solution is not available for this question yet.
Question 13 MSQ · 3.0 marks
Consider a transformer model using scaled dot-product attention:
[[IMAGE:0bc72138559e0d4e_8_46]]
Assume a fixed input size and embedding dimension [[IMAGE:0bc72138559e0d4e_8_47]] . The model may use either single-head or
multi-head attention. Which of the following statements are correct?


Scaling by [[IMAGE:0bc72138559e0d4e_8_48]] helps prevent the dot-product values from becoming too large
when [[IMAGE:0bc72138559e0d4e_8_49]] is high.


In multi-head attention, all heads share the same learned projections of
[[IMAGE:0bc72138559e0d4e_8_50]] .

If all attention heads learn identical projection matrices, multi-head attention
behaves equivalently to single-head attention.
For a fixed embedding dimension, increasing the number of heads always
increases the total number of parameters in the query, key, and value projections.
A published solution is not available for this question yet.
Question 14 NAT · 3.0 marks
A text corpus contains the following three sentences:
1. students read the books in library
2. teachers read the books after class
3. students and teachers discuss in library
After removing the stop words "the'', "in", and "and'', the vocabulary becomes:
[[IMAGE:0bc72138559e0d4e_9_51]]
Assume the context window size is [[IMAGE:0bc72138559e0d4e_9_52]] (one word to the left and one word to the right).
Based on the above data, answer the given subquestions.
Construct the co-occurrence matrix [[IMAGE:0bc72138559e0d4e_9_53]] , where rows represent target words and columns represent
context words. Let [[IMAGE:0bc72138559e0d4e_9_54]] denote the element in the [[IMAGE:0bc72138559e0d4e_9_55]] -th row and [[IMAGE:0bc72138559e0d4e_9_56]] -th column. Enter the value of
[[IMAGE:0bc72138559e0d4e_9_57]] .







A published solution is not available for this question yet.
Question 15 MCQ · 3.0 marks
A text corpus contains the following three sentences:
1. students read the books in library
2. teachers read the books after class
3. students and teachers discuss in library
After removing the stop words "the'', "in", and "and'', the vocabulary becomes:
[[IMAGE:0bc72138559e0d4e_9_51]]
Assume the context window size is [[IMAGE:0bc72138559e0d4e_9_52]] (one word to the left and one word to the right).
Based on the above data, answer the given subquestions.
Based on the co-occurrence matrix constructed previously, determine whether [[IMAGE:0bc72138559e0d4e_10_58]]
is positive, zero, or negative. (Use natural log).



Positive
Zero
Negative
A published solution is not available for this question yet.
Question 16 MSQ · 3.0 marks
A text corpus contains the following three sentences:
1. students read the books in library
2. teachers read the books after class
3. students and teachers discuss in library
After removing the stop words "the'', "in", and "and'', the vocabulary becomes:
[[IMAGE:0bc72138559e0d4e_9_51]]
Assume the context window size is [[IMAGE:0bc72138559e0d4e_9_52]] (one word to the left and one word to the right).
Based on the above data, answer the given subquestions.
Which of the following problems of the co-occurrence matrix are addressed by applying Singular
Value Decomposition (SVD)? Choose all the correct options.


High dimensional representation
Sparse matrix with many zeros
Captures latent semantic similarity between words
Increases vocabulary size
Produces dense low-dimensional embeddings
A published solution is not available for this question yet.
Question 17 NAT · 2.0 marks
You are given a transformer attention setup with the following configuration:
• Sequence Length : [[IMAGE:0bc72138559e0d4e_10_59]]
• Number of Heads : [[IMAGE:0bc72138559e0d4e_10_60]]
• Embedding dimension : [[IMAGE:0bc72138559e0d4e_10_61]]
• Input [[IMAGE:0bc72138559e0d4e_10_62]]
•
[[IMAGE:0bc72138559e0d4e_11_63]]
• [[IMAGE:0bc72138559e0d4e_11_64]]
• [[IMAGE:0bc72138559e0d4e_11_65]]
• [[IMAGE:0bc72138559e0d4e_11_66]]
• [[IMAGE:0bc72138559e0d4e_11_67]]
Based on the above data, answer the given subquestions.
Suppose [[IMAGE:0bc72138559e0d4e_11_68]] , [[IMAGE:0bc72138559e0d4e_11_69]] , [[IMAGE:0bc72138559e0d4e_11_70]] and [[IMAGE:0bc72138559e0d4e_11_71]] . The scaled dot-product attention operation for
a single head, given by:
[[IMAGE:0bc72138559e0d4e_11_72]]
Compute the total number of elements in the resulting attention output.














A published solution is not available for this question yet.
Question 18 NAT · 2.0 marks
You are given a transformer attention setup with the following configuration:
• Sequence Length : [[IMAGE:0bc72138559e0d4e_10_59]]
• Number of Heads : [[IMAGE:0bc72138559e0d4e_10_60]]
• Embedding dimension : [[IMAGE:0bc72138559e0d4e_10_61]]
• Input [[IMAGE:0bc72138559e0d4e_10_62]]
•
[[IMAGE:0bc72138559e0d4e_11_63]]
• [[IMAGE:0bc72138559e0d4e_11_64]]
• [[IMAGE:0bc72138559e0d4e_11_65]]
• [[IMAGE:0bc72138559e0d4e_11_66]]
• [[IMAGE:0bc72138559e0d4e_11_67]]
Based on the above data, answer the given subquestions.
Suppose [[IMAGE:0bc72138559e0d4e_11_73]] , [[IMAGE:0bc72138559e0d4e_11_74]] , [[IMAGE:0bc72138559e0d4e_11_75]] and [[IMAGE:0bc72138559e0d4e_11_76]] . What will be the number of parameters in the
Multihead Attention ? (ignore the bias)













A published solution is not available for this question yet.
Question 19 NAT · 2.0 marks
You are given a transformer attention setup with the following configuration:
• Sequence Length : [[IMAGE:0bc72138559e0d4e_10_59]]
• Number of Heads : [[IMAGE:0bc72138559e0d4e_10_60]]
• Embedding dimension : [[IMAGE:0bc72138559e0d4e_10_61]]
• Input [[IMAGE:0bc72138559e0d4e_10_62]]
•
[[IMAGE:0bc72138559e0d4e_11_63]]
• [[IMAGE:0bc72138559e0d4e_11_64]]
• [[IMAGE:0bc72138559e0d4e_11_65]]
• [[IMAGE:0bc72138559e0d4e_11_66]]
• [[IMAGE:0bc72138559e0d4e_11_67]]
Based on the above data, answer the given subquestions.
Assume a Feed-Forward Network (FFN) follows the Multi-Head Attention layer in the encoder. The
FFN consists of two linear transformations:
[[IMAGE:0bc72138559e0d4e_12_77]]
Calculate the total number of parameters in the FFN layer (including bias terms), where [[IMAGE:0bc72138559e0d4e_12_78]]
and [[IMAGE:0bc72138559e0d4e_12_79]] .












A published solution is not available for this question yet.
Question 20 NAT · 2.0 marks
Assume that your CBOW model outputs a probability distribution over a vocabulary of 20,000
words for a given context. If the correct target word is word number 150, and the model's
predicted probability for this word is 0.2, calculate the cross-entropy loss for this prediction. Enter
the answer correct to two decimal places. (use natural log)
A published solution is not available for this question yet.
Question 21 NAT · 3.0 marks
In the Skip-gram model with a window size of 2 (on each side), how many **unique** pairs of target
and context words will be generated for the following sentence.
Note: Ignore punctuation in your calculation
'**Learning from data helps build intelligent systems for society**'
A published solution is not available for this question yet.