cs3004_2026T2_ET_FN.pdf
Deep Learning · End Term · May 2026 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 NAT · 2.0 marks
A neuron takes an input vector [[IMAGE:8b451edb9ea00ccd_2_2]] , weight vector [[IMAGE:8b451edb9ea00ccd_2_3]] , and bias
[[IMAGE:8b451edb9ea00ccd_2_4]] . The neuron uses Parametric ReLU activation function parameterized by [[IMAGE:8b451edb9ea00ccd_2_5]] :
[[IMAGE:8b451edb9ea00ccd_2_6]]
Calculate the exact activation output [[IMAGE:8b451edb9ea00ccd_2_7]] for this neuron. Enter the answer correct to one decimal
place.






A published solution is not available for this question yet.
Question 3 NAT · 2.0 marks
A satellite image of size [[IMAGE:8b451edb9ea00ccd_2_8]] is given as input to a convolutional layer. Suppose the
layer uses [[IMAGE:8b451edb9ea00ccd_2_9]] kernels, each of size [[IMAGE:8b451edb9ea00ccd_2_10]] , and produces an output of size [[IMAGE:8b451edb9ea00ccd_2_11]] .
Assume zero padding (
[[IMAGE:8b451edb9ea00ccd_3_12]] ). What is the value of the stride [[IMAGE:8b451edb9ea00ccd_3_13]] ?






A published solution is not available for this question yet.
Question 4 NAT · 3.0 marks
Suppose we want to minimize the following function:
[[IMAGE:8b451edb9ea00ccd_3_14]]
Let the initial point [[IMAGE:8b451edb9ea00ccd_3_15]] .
Based on the above data, answer the given subquestions.
Using vanilla gradient descent with [[IMAGE:8b451edb9ea00ccd_3_16]] , compute the iterate [[IMAGE:8b451edb9ea00ccd_3_17]] after one update, and
enter your answer as [[IMAGE:8b451edb9ea00ccd_3_18]] . Enter the answer correct to one decimal place.





A published solution is not available for this question yet.
Question 5 MCQ · 2.0 marks
Suppose we want to minimize the following function:
[[IMAGE:8b451edb9ea00ccd_3_14]]
Let the initial point [[IMAGE:8b451edb9ea00ccd_3_15]] .
Based on the above data, answer the given subquestions.
Along which coordinate does gradient descent with [[IMAGE:8b451edb9ea00ccd_4_19]] fail to converge?



[[IMAGE:8b451edb9ea00ccd_4_20]] -coordinate

[[IMAGE:8b451edb9ea00ccd_4_21]] -coordinate

Both coordinates diverge equally.
Neither; both converge at the same rate.
A published solution is not available for this question yet.
Question 6 NAT · 3.0 marks
Suppose we want to minimize the following function:
[[IMAGE:8b451edb9ea00ccd_3_14]]
Let the initial point [[IMAGE:8b451edb9ea00ccd_3_15]] .
Based on the above data, answer the given subquestions.
[[IMAGE:8b451edb9ea00ccd_4_22]]



A published solution is not available for this question yet.
Question 7 MCQ · 2.0 marks
Suppose we want to minimize the following function:
[[IMAGE:8b451edb9ea00ccd_3_14]]
Let the initial point [[IMAGE:8b451edb9ea00ccd_3_15]] .
Based on the above data, answer the given subquestions.
How does AdaGrad address the behavior observed with vanilla gradient descent?


The learning rate is reduced 25 times from the initial learning rate 0.1 along
the [[IMAGE:8b451edb9ea00ccd_5_23]] -direction.

The learning rate is reduced 25 times from the initial learning rate 0.1 along
the [[IMAGE:8b451edb9ea00ccd_5_24]] -direction.

The learning rate is reduced 5 times from the initial learning rate 0.1 along the
[[IMAGE:8b451edb9ea00ccd_5_25]] -direction.

The learning rate is reduced [[IMAGE:8b451edb9ea00ccd_5_26]] times from the initial learning rate [[IMAGE:8b451edb9ea00ccd_5_27]] along
both the [[IMAGE:8b451edb9ea00ccd_5_28]] and [[IMAGE:8b451edb9ea00ccd_5_29]] directions.




A published solution is not available for this question yet.
Question 8 MSQ · 2.0 marks
Consider a feedforward neural network with [[IMAGE:8b451edb9ea00ccd_5_30]] hidden layers. The pre-activation at layer [[IMAGE:8b451edb9ea00ccd_5_31]] is
[[IMAGE:8b451edb9ea00ccd_5_32]] and the activation is [[IMAGE:8b451edb9ea00ccd_5_33]] , where [[IMAGE:8b451edb9ea00ccd_5_34]] is the activation function. Suppose
every hidden-layer activation is the identity function [[IMAGE:8b451edb9ea00ccd_5_35]] , and the output activation is linear.
Which of the following statements are correct?






The entire network computes a function of the form [[IMAGE:8b451edb9ea00ccd_5_36]] for some
single matrix [[IMAGE:8b451edb9ea00ccd_5_37]] and vector [[IMAGE:8b451edb9ea00ccd_5_38]] , regardless of the number of hidden layers.



Adding more hidden layers with identity activations increases the
representational capacity beyond that of a single-layer linear model.
Replacing the identity activation with a nonlinear function such as sigmoid
enables the network to learn nonlinear decision boundaries.
With identity activations, the network can still learn arbitrary decision
boundaries provided enough hidden neurons are used.
A published solution is not available for this question yet.
Question 9 NAT · 3.0 marks
A multi-head attention block uses [[IMAGE:8b451edb9ea00ccd_6_39]] , [[IMAGE:8b451edb9ea00ccd_6_40]] heads, and [[IMAGE:8b451edb9ea00ccd_6_41]] . Counting only the
weight matrices [[IMAGE:8b451edb9ea00ccd_6_42]] (each of dimension [[IMAGE:8b451edb9ea00ccd_6_43]] ) for every head, plus the output
projection [[IMAGE:8b451edb9ea00ccd_6_44]] (dimension [[IMAGE:8b451edb9ea00ccd_6_45]] ), and ignoring all biases, compute the total number of
parameters in this multi-head attention block.







A published solution is not available for this question yet.
Question 10 MCQ · 3.0 marks
In a self-attention mechanism, each input vector has dimension [[IMAGE:8b451edb9ea00ccd_6_46]] . The queries, keys, and values
are computed using separate linear transformations, each mapping from [[IMAGE:8b451edb9ea00ccd_6_47]] . Suppose
there are [[IMAGE:8b451edb9ea00ccd_6_48]] input vectors.
Based on the above data, answer the given subquestions.
How many learnable parameters (including both weights and biases) are used to compute the
queries, keys, and values?



[[IMAGE:8b451edb9ea00ccd_7_49]]

[[IMAGE:8b451edb9ea00ccd_7_50]]

[[IMAGE:8b451edb9ea00ccd_7_51]]

[[IMAGE:8b451edb9ea00ccd_7_52]]

A published solution is not available for this question yet.
Question 11 NAT · 2.0 marks
In a self-attention mechanism, each input vector has dimension [[IMAGE:8b451edb9ea00ccd_6_46]] . The queries, keys, and values
are computed using separate linear transformations, each mapping from [[IMAGE:8b451edb9ea00ccd_6_47]] . Suppose
there are [[IMAGE:8b451edb9ea00ccd_6_48]] input vectors.
Based on the above data, answer the given subquestions.
How many attention weights [[IMAGE:8b451edb9ea00ccd_7_53]] are computed in total if we have [[IMAGE:8b451edb9ea00ccd_7_54]] and [[IMAGE:8b451edb9ea00ccd_7_55]] ?






A published solution is not available for this question yet.
Question 12 MCQ · 2.0 marks
In a self-attention mechanism, each input vector has dimension [[IMAGE:8b451edb9ea00ccd_6_46]] . The queries, keys, and values
are computed using separate linear transformations, each mapping from [[IMAGE:8b451edb9ea00ccd_6_47]] . Suppose
there are [[IMAGE:8b451edb9ea00ccd_6_48]] input vectors.
Based on the above data, answer the given subquestions.
How many learnable parameters (including both weights and biases) are there in a fully connected
shallow network that maps all [[IMAGE:8b451edb9ea00ccd_7_56]] inputs to all [[IMAGE:8b451edb9ea00ccd_7_57]] outputs?





[[IMAGE:8b451edb9ea00ccd_7_58]]

[[IMAGE:8b451edb9ea00ccd_7_59]]

[[IMAGE:8b451edb9ea00ccd_7_60]]

[[IMAGE:8b451edb9ea00ccd_7_61]]

A published solution is not available for this question yet.
Question 13 NAT · 2.0 marks
[[IMAGE:8b451edb9ea00ccd_8_62]]
Based on the above data, answer the given subquestions.
Which token receives the lowest dot-product score [[IMAGE:8b451edb9ea00ccd_8_63]] from query [[IMAGE:8b451edb9ea00ccd_8_64]] ? Enter the token index (1,
2, or 3).



A published solution is not available for this question yet.
Question 14 NAT · 3.0 marks
[[IMAGE:8b451edb9ea00ccd_8_62]]
Based on the above data, answer the given subquestions.
Compute the attention weight [[IMAGE:8b451edb9ea00ccd_9_65]] (the weight that query [[IMAGE:8b451edb9ea00ccd_9_66]] places on token 3). Enter the answer
correct to three decimal places.



A published solution is not available for this question yet.
Question 15 NAT · 3.0 marks
[[IMAGE:8b451edb9ea00ccd_8_62]]
Based on the above data, answer the given subquestions.
[[IMAGE:8b451edb9ea00ccd_9_67]]


A published solution is not available for this question yet.
Question 16 MSQ · 3.0 marks
In a vanilla (non-attention) encoder-decoder model, the encoder compresses the entire input
sequence into a single fixed-dimensional vector [[IMAGE:8b451edb9ea00ccd_10_68]] , such as [[IMAGE:8b451edb9ea00ccd_10_69]] , where [[IMAGE:8b451edb9ea00ccd_10_70]] is the final hidden
state of an RNN encoder. The decoder uses this same vector to generate every output token. In
practice, the translation quality of such a model tends to degrade as the length of the source
sentence increases. Which of the following best explains this degradation? Select all that apply



The decoder's softmax layer has a fixed vocabulary size and therefore cannot
generate longer sentences.
The fixed-dimensional vector [[IMAGE:8b451edb9ea00ccd_10_71]] becomes an information bottleneck as the
source sentence length increases.

Teacher forcing is valid only for source sentences shorter than the maximum
training length.
Longer source sentences necessarily require a higher-dimensional one-hot
representation, which increases the input noise.
A published solution is not available for this question yet.
Question 17 NAT · 3.0 marks
Consider a Skip-gram model with negative sampling. For a positive word-context pair [[IMAGE:8b451edb9ea00ccd_10_72]] ,
suppose the loss for one training example is
[[IMAGE:8b451edb9ea00ccd_10_73]]
where [[IMAGE:8b451edb9ea00ccd_10_74]] is a logistic function. Suppose the initial vectors are
[[IMAGE:8b451edb9ea00ccd_10_75]]
The word vector w is updated using the gradient descent algorithm using a learning rate of
[[IMAGE:8b451edb9ea00ccd_10_76]]
Based on the above data, answer the given subquestions.
Calculate the loss [[IMAGE:8b451edb9ea00ccd_11_77]] for the given positive word-context pair [[IMAGE:8b451edb9ea00ccd_11_78]] and negative sample. Give your
answer correct to three decimal places.







A published solution is not available for this question yet.
Question 18 NAT · 3.0 marks
Consider a Skip-gram model with negative sampling. For a positive word-context pair [[IMAGE:8b451edb9ea00ccd_10_72]] ,
suppose the loss for one training example is
[[IMAGE:8b451edb9ea00ccd_10_73]]
where [[IMAGE:8b451edb9ea00ccd_10_74]] is a logistic function. Suppose the initial vectors are
[[IMAGE:8b451edb9ea00ccd_10_75]]
The word vector w is updated using the gradient descent algorithm using a learning rate of
[[IMAGE:8b451edb9ea00ccd_10_76]]
Based on the above data, answer the given subquestions.
The word vector w is updated using the gradient descent algorithm. Using gradient descent,
obtain the updated word vector [[IMAGE:8b451edb9ea00ccd_11_79]] , and calculate the cosine similarity between [[IMAGE:8b451edb9ea00ccd_11_80]] and [[IMAGE:8b451edb9ea00ccd_11_81]] .
Enter the answer correct to three decimal places.








A published solution is not available for this question yet.
Question 19 NAT · 3.0 marks
Consider a corpus containing [[IMAGE:8b451edb9ea00ccd_11_82]] total words. Based on this, answer the given
subquestions:
A word [[IMAGE:8b451edb9ea00ccd_12_83]] occurs [[IMAGE:8b451edb9ea00ccd_12_84]] times, a context [[IMAGE:8b451edb9ea00ccd_12_85]] occurs [[IMAGE:8b451edb9ea00ccd_12_86]] times, and the pair
[[IMAGE:8b451edb9ea00ccd_12_87]] co-occurs [[IMAGE:8b451edb9ea00ccd_12_88]] times. Compute [[IMAGE:8b451edb9ea00ccd_12_89]] . Use natural log and enter the
answer correct to two decimal places.








A published solution is not available for this question yet.
Question 20 NAT · 1.0 marks
Consider a corpus containing [[IMAGE:8b451edb9ea00ccd_11_82]] total words. Based on this, answer the given
subquestions:
A second pair [[IMAGE:8b451edb9ea00ccd_12_90]] has
[[IMAGE:8b451edb9ea00ccd_12_91]]
Suppose its PMI is [[IMAGE:8b451edb9ea00ccd_12_92]] . Under the PPMI (positive PMI) transformation, what value is stored
for this pair?




A published solution is not available for this question yet.
Question 21 MCQ · 3.0 marks
A neural network is trained using mini-batch gradient descent. Suppose the mini-batch size is
increased from [[IMAGE:8b451edb9ea00ccd_13_93]] to [[IMAGE:8b451edb9ea00ccd_13_94]] , while the learning rate and training dataset remain unchanged. Which
of the following statements is correct?


Each update uses a gradient estimate with higher variance, while the number
of updates per epoch decreases.
Each update uses a gradient estimate with lower variance, while the number
of updates per epoch decreases.
The number of parameter updates per epoch increases by a factor of four,
leading to faster convergence.
The larger batch size guarantees convergence to a better local minimum.
A published solution is not available for this question yet.