cs3004_2025T3_ET_FN.pdf
Deep Learning · End Term · Sep 2025 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 2.0 marks
What is the primary advantage of word embeddings compared to one-hot encoding?
Simpler implementation
Captures semantic relationships
They automatically correct spelling errors
Word embeddings are sparser than one-hot encoding
A published solution is not available for this question yet.
Question 3 MCQ · 2.0 marks
What is the primary advantage of Transformer encoders over RNN encoders for long sequences?
Lower memory usage
Parallelization during training
Smaller model size
Simpler architecture
A published solution is not available for this question yet.
Question 4 MSQ · 3.0 marks
Suppose that we implement an OR function with two boolean inputs using the McCulloch Pitts
(MP) neuron. Assume that the neuron does not have any inhibitory input. Which of the following
option(s) represents correct [[IMAGE:64b1fdfafe1feffe_3_2]] and the Number of Correctly Classified (NCC) data points for various
values of threshold [[IMAGE:64b1fdfafe1feffe_3_3]] .
[[IMAGE:64b1fdfafe1feffe_3_4]]



[[IMAGE:64b1fdfafe1feffe_3_5]] , NCC = 4

[[IMAGE:64b1fdfafe1feffe_3_6]] , NCC = 4

[[IMAGE:64b1fdfafe1feffe_3_7]] , NCC = 3

[[IMAGE:64b1fdfafe1feffe_3_8]] , NCC = 1

[[IMAGE:64b1fdfafe1feffe_3_9]] , NCC = 2

[[IMAGE:64b1fdfafe1feffe_3_10]] , NCC = 0

A published solution is not available for this question yet.
Question 5 MSQ · 2.0 marks
Which of the following optimizers use momentum (or momentum-like mechanisms) in updating
gradients?
NAG
AdaGrad
RMSProp
AdaDelta
ADAM
A published solution is not available for this question yet.
Question 6 MSQ · 2.0 marks
Which of the following statements are true for the saturated neurons problem? Select all that
apply.
ReLU activation function is preferred over tanh, as it is less likely to saturate.
Tanh is preferred over sigmoid as it solves the problem of neuron saturation.
Leaky ReLU mitigates the problem of neuron saturation compared to ReLU.
If a neuron saturates in a network, the weights do not get updated further.
A published solution is not available for this question yet.
Question 7 MSQ · 3.0 marks
You are training a deep feedforward neural network (100 layers) for binary classification with a
sigmoid output with tanh activation in some hidden layers and ReLU activations in some hidden
layers. You notice that gradients in some layers vanish early, preventing those weights from
updating, even though the network hasn’t converged.
Which of the following fixes could help?
Increase the size of your training set
Replace ReLU activations with leaky ReLUs
Use appropriate weight initialization
A published solution is not available for this question yet.
Question 8 NAT · 3.0 marks
A transformer with a single self-attention layer has batch size [[IMAGE:64b1fdfafe1feffe_4_11]] , context length (sequence
length) [[IMAGE:64b1fdfafe1feffe_4_12]] , and number of heads [[IMAGE:64b1fdfafe1feffe_4_13]] . How many attention scores are computed in total
across two heads and all sequences for one forward pass?



A published solution is not available for this question yet.
Question 9 NAT · 3.0 marks
[[IMAGE:64b1fdfafe1feffe_5_14]]

A published solution is not available for this question yet.
Question 10 NAT · 3.0 marks
Given the input matrix [[IMAGE:64b1fdfafe1feffe_6_15]] and kernel [[IMAGE:64b1fdfafe1feffe_6_16]] :
[[IMAGE:64b1fdfafe1feffe_6_17]]
• Perform convolution of [[IMAGE:64b1fdfafe1feffe_6_18]] over [[IMAGE:64b1fdfafe1feffe_6_19]] with stride [[IMAGE:64b1fdfafe1feffe_6_20]] and no padding to obtain matrix [[IMAGE:64b1fdfafe1feffe_6_21]] .
• Apply average pooling over [[IMAGE:64b1fdfafe1feffe_6_22]] of size [[IMAGE:64b1fdfafe1feffe_6_23]] with stride [[IMAGE:64b1fdfafe1feffe_6_24]] and no padding to obtain [[IMAGE:64b1fdfafe1feffe_6_25]] .
• Apply the **ReLU activation** on [[IMAGE:64b1fdfafe1feffe_6_26]] to obtain the final output [[IMAGE:64b1fdfafe1feffe_6_27]] .
Given that
[[IMAGE:64b1fdfafe1feffe_6_28]]
compute
[[IMAGE:64b1fdfafe1feffe_6_29]]
where [[IMAGE:64b1fdfafe1feffe_6_30]] is the **center element** of the kernel.
**Submit your answer correct to two decimal places.**
















A published solution is not available for this question yet.
Question 11 NAT · 2.0 marks
You are designing a Skip-Gram model to learn word embeddings. In this iteration, the model
focuses on a specific **center word** and attempts to predict the surrounding **context words** within
a window size of [[IMAGE:64b1fdfafe1feffe_7_31]] .
The sentence is: **Deep neural networks learn complex data patterns.**
We are currently processing the center word: **networks**.
Vocabulary = {0: complex, 1: data, 2: deep, 3: learn, 4: models, 5: networks, 6: neural, 7: patterns }
The one-hot vector for the center word ( [[IMAGE:64b1fdfafe1feffe_7_32]] ) is a column vector of size [[IMAGE:64b1fdfafe1feffe_7_33]] .
[[IMAGE:64b1fdfafe1feffe_7_34]]
The model's weight matrices are given below:
**Input/Center Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_35]] **):**
[[IMAGE:64b1fdfafe1feffe_7_36]]
(Note: Row 5 corresponds to word index 5 `networks')
**Output/Context Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_37]] **):**
[[IMAGE:64b1fdfafe1feffe_7_38]]
For the center word [[IMAGE:64b1fdfafe1feffe_7_39]] , the hidden layer representation is given by:
[[IMAGE:64b1fdfafe1feffe_7_40]]
(Or simply selecting the row corresponding to the center word).
The unnormalized score vector [[IMAGE:64b1fdfafe1feffe_7_41]] for all words in the vocabulary is calculated as:
[[IMAGE:64b1fdfafe1feffe_7_42]]
Based on the above data, answer the given subquestions.
In the Skip-gram model with a window size of [[IMAGE:64b1fdfafe1feffe_8_43]] (on each side), how many **unique** input and
output pairs will be generated for the following sentence. The sentence is: **deep neural networks**
**learn complex data patterns.**













A published solution is not available for this question yet.
Question 12 MCQ · 1.0 marks
You are designing a Skip-Gram model to learn word embeddings. In this iteration, the model
focuses on a specific **center word** and attempts to predict the surrounding **context words** within
a window size of [[IMAGE:64b1fdfafe1feffe_7_31]] .
The sentence is: **Deep neural networks learn complex data patterns.**
We are currently processing the center word: **networks**.
Vocabulary = {0: complex, 1: data, 2: deep, 3: learn, 4: models, 5: networks, 6: neural, 7: patterns }
The one-hot vector for the center word ( [[IMAGE:64b1fdfafe1feffe_7_32]] ) is a column vector of size [[IMAGE:64b1fdfafe1feffe_7_33]] .
[[IMAGE:64b1fdfafe1feffe_7_34]]
The model's weight matrices are given below:
**Input/Center Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_35]] **):**
[[IMAGE:64b1fdfafe1feffe_7_36]]
(Note: Row 5 corresponds to word index 5 `networks')
**Output/Context Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_37]] **):**
[[IMAGE:64b1fdfafe1feffe_7_38]]
For the center word [[IMAGE:64b1fdfafe1feffe_7_39]] , the hidden layer representation is given by:
[[IMAGE:64b1fdfafe1feffe_7_40]]
(Or simply selecting the row corresponding to the center word).
The unnormalized score vector [[IMAGE:64b1fdfafe1feffe_7_41]] for all words in the vocabulary is calculated as:
[[IMAGE:64b1fdfafe1feffe_7_42]]
Based on the above data, answer the given subquestions.
Calculate the hidden layer representation vector [[IMAGE:64b1fdfafe1feffe_8_44]] for the center word "networks".













[[IMAGE:64b1fdfafe1feffe_8_45]]

[[IMAGE:64b1fdfafe1feffe_8_46]]

[[IMAGE:64b1fdfafe1feffe_8_47]]

[[IMAGE:64b1fdfafe1feffe_8_48]]

A published solution is not available for this question yet.
Question 13 NAT · 2.0 marks
You are designing a Skip-Gram model to learn word embeddings. In this iteration, the model
focuses on a specific **center word** and attempts to predict the surrounding **context words** within
a window size of [[IMAGE:64b1fdfafe1feffe_7_31]] .
The sentence is: **Deep neural networks learn complex data patterns.**
We are currently processing the center word: **networks**.
Vocabulary = {0: complex, 1: data, 2: deep, 3: learn, 4: models, 5: networks, 6: neural, 7: patterns }
The one-hot vector for the center word ( [[IMAGE:64b1fdfafe1feffe_7_32]] ) is a column vector of size [[IMAGE:64b1fdfafe1feffe_7_33]] .
[[IMAGE:64b1fdfafe1feffe_7_34]]
The model's weight matrices are given below:
**Input/Center Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_35]] **):**
[[IMAGE:64b1fdfafe1feffe_7_36]]
(Note: Row 5 corresponds to word index 5 `networks')
**Output/Context Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_37]] **):**
[[IMAGE:64b1fdfafe1feffe_7_38]]
For the center word [[IMAGE:64b1fdfafe1feffe_7_39]] , the hidden layer representation is given by:
[[IMAGE:64b1fdfafe1feffe_7_40]]
(Or simply selecting the row corresponding to the center word).
The unnormalized score vector [[IMAGE:64b1fdfafe1feffe_7_41]] for all words in the vocabulary is calculated as:
[[IMAGE:64b1fdfafe1feffe_7_42]]
Based on the above data, answer the given subquestions.
We want to predict the immediate **next word** to the word "networks" in the sentence, which is
"learn". Calculate the unnormalized score ( [[IMAGE:64b1fdfafe1feffe_8_49]] ) for this specific target word.













A published solution is not available for this question yet.
Question 14 NAT · 2.0 marks
You are designing a Skip-Gram model to learn word embeddings. In this iteration, the model
focuses on a specific **center word** and attempts to predict the surrounding **context words** within
a window size of [[IMAGE:64b1fdfafe1feffe_7_31]] .
The sentence is: **Deep neural networks learn complex data patterns.**
We are currently processing the center word: **networks**.
Vocabulary = {0: complex, 1: data, 2: deep, 3: learn, 4: models, 5: networks, 6: neural, 7: patterns }
The one-hot vector for the center word ( [[IMAGE:64b1fdfafe1feffe_7_32]] ) is a column vector of size [[IMAGE:64b1fdfafe1feffe_7_33]] .
[[IMAGE:64b1fdfafe1feffe_7_34]]
The model's weight matrices are given below:
**Input/Center Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_35]] **):**
[[IMAGE:64b1fdfafe1feffe_7_36]]
(Note: Row 5 corresponds to word index 5 `networks')
**Output/Context Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_37]] **):**
[[IMAGE:64b1fdfafe1feffe_7_38]]
For the center word [[IMAGE:64b1fdfafe1feffe_7_39]] , the hidden layer representation is given by:
[[IMAGE:64b1fdfafe1feffe_7_40]]
(Or simply selecting the row corresponding to the center word).
The unnormalized score vector [[IMAGE:64b1fdfafe1feffe_7_41]] for all words in the vocabulary is calculated as:
[[IMAGE:64b1fdfafe1feffe_7_42]]
Based on the above data, answer the given subquestions.
Identify the word in the vocabulary that has the highest probability of being predicted as a context
word given the center word 'networks'. Provide the corresponding **integer index (key)** from the
vocabulary list as your answer.












A published solution is not available for this question yet.
Question 15 NAT · 2.0 marks
You are designing a Skip-Gram model to learn word embeddings. In this iteration, the model
focuses on a specific **center word** and attempts to predict the surrounding **context words** within
a window size of [[IMAGE:64b1fdfafe1feffe_7_31]] .
The sentence is: **Deep neural networks learn complex data patterns.**
We are currently processing the center word: **networks**.
Vocabulary = {0: complex, 1: data, 2: deep, 3: learn, 4: models, 5: networks, 6: neural, 7: patterns }
The one-hot vector for the center word ( [[IMAGE:64b1fdfafe1feffe_7_32]] ) is a column vector of size [[IMAGE:64b1fdfafe1feffe_7_33]] .
[[IMAGE:64b1fdfafe1feffe_7_34]]
The model's weight matrices are given below:
**Input/Center Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_35]] **):**
[[IMAGE:64b1fdfafe1feffe_7_36]]
(Note: Row 5 corresponds to word index 5 `networks')
**Output/Context Weight Matrix (** [[IMAGE:64b1fdfafe1feffe_7_37]] **):**
[[IMAGE:64b1fdfafe1feffe_7_38]]
For the center word [[IMAGE:64b1fdfafe1feffe_7_39]] , the hidden layer representation is given by:
[[IMAGE:64b1fdfafe1feffe_7_40]]
(Or simply selecting the row corresponding to the center word).
The unnormalized score vector [[IMAGE:64b1fdfafe1feffe_7_41]] for all words in the vocabulary is calculated as:
[[IMAGE:64b1fdfafe1feffe_7_42]]
Based on the above data, answer the given subquestions.
How many total weights from Win and Wout combined will be updated?












A published solution is not available for this question yet.
Question 16 NAT · 2.0 marks
Suppose you are given three encoder hidden states at time [[IMAGE:64b1fdfafe1feffe_10_50]] :
[[IMAGE:64b1fdfafe1feffe_10_51]]
The previous decoder hidden state is:
[[IMAGE:64b1fdfafe1feffe_10_52]]
Given the attention score function:
[[IMAGE:64b1fdfafe1feffe_10_53]]
where the hyperbolic tangent function is defined as:
[[IMAGE:64b1fdfafe1feffe_10_54]]
and
[[IMAGE:64b1fdfafe1feffe_10_55]]
Based on the above data, answer the given subquestions.
Compute the attention score for hidden state [[IMAGE:64b1fdfafe1feffe_10_56]] using the given function. Submit the final answer
correct to two decimal places.







A published solution is not available for this question yet.
Question 17 NAT · 4.0 marks
Suppose you are given three encoder hidden states at time [[IMAGE:64b1fdfafe1feffe_10_50]] :
[[IMAGE:64b1fdfafe1feffe_10_51]]
The previous decoder hidden state is:
[[IMAGE:64b1fdfafe1feffe_10_52]]
Given the attention score function:
[[IMAGE:64b1fdfafe1feffe_10_53]]
where the hyperbolic tangent function is defined as:
[[IMAGE:64b1fdfafe1feffe_10_54]]
and
[[IMAGE:64b1fdfafe1feffe_10_55]]
Based on the above data, answer the given subquestions.
Normalize the attention scores using the softmax function to obtain the attention weights [[IMAGE:64b1fdfafe1feffe_11_57]] .
Submit [[IMAGE:64b1fdfafe1feffe_11_58]] (i.e., first element of the [[IMAGE:64b1fdfafe1feffe_11_59]] vector). Submit the final answer correct to two decimal
places.
[[IMAGE:64b1fdfafe1feffe_11_60]]










A published solution is not available for this question yet.
Question 18 NAT · 2.0 marks
Suppose you are given three encoder hidden states at time [[IMAGE:64b1fdfafe1feffe_10_50]] :
[[IMAGE:64b1fdfafe1feffe_10_51]]
The previous decoder hidden state is:
[[IMAGE:64b1fdfafe1feffe_10_52]]
Given the attention score function:
[[IMAGE:64b1fdfafe1feffe_10_53]]
where the hyperbolic tangent function is defined as:
[[IMAGE:64b1fdfafe1feffe_10_54]]
and
[[IMAGE:64b1fdfafe1feffe_10_55]]
Based on the above data, answer the given subquestions.
Calculate the context vector [[IMAGE:64b1fdfafe1feffe_11_61]] :
[[IMAGE:64b1fdfafe1feffe_11_62]]
Provide the sum of all the elements of [[IMAGE:64b1fdfafe1feffe_11_63]] . Submit the final answer correct to two decimal places.









A published solution is not available for this question yet.