da2001_2026T2_Q1_NA.pdf
Introduction to Deep Learning and Generative AI(DL GENAI) · Quiz 1 · May 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 NAT · 3.0 marks
Consider the following data point and weight vector in a simple single layer neural network.
(Assume that there is no bias.)
x = [1,0]
w = [0,1]
Once the data flows through the simple network and is passed through the activation function,
the final output comes out to be 0.5.
Once you have identified the activation function, enter the maximum value that the derivative of
the activation function can take.
A published solution is not available for this question yet.
Question 3 NAT · 3.0 marks
A BatchNorm layer is applied to the output of a convolutional layer. The output feature map has
dimensions:
[[IMAGE:193adccb815e9fcb_3_2]]
where:
• 32 = Batch size (number of images in the mini-batch)
• 28 = Height of each feature map
• 28 = Width of each feature map
• 64 = Number of channels
How many learnable parameters does this BatchNorm layer contain?

A published solution is not available for this question yet.
Question 4 MCQ · 3.0 marks
Consider a McCulloch-Pitts (MP) neuron with binary inputs, unit positive weights, and no inhibitory
connections.
Given the following data points:
a = [0,0]
b = [1,1]
c = [1,0]
The neuron must satisfy:
a → 0
b → 0
c → 1
Which of the following statements is TRUE (θ is the threshold for the neuron to fire.)?
This is possible with [[IMAGE:193adccb815e9fcb_4_3]] = 1

This is possible with [[IMAGE:193adccb815e9fcb_4_4]] = 2

This is possible with some threshold value
This is impossible for a standard MP neuron
A published solution is not available for this question yet.
Question 5 MCQ · 3.0 marks
Consider the following pytorch code:
[[IMAGE:193adccb815e9fcb_4_5]]
What is the output of the above code?

[[IMAGE:193adccb815e9fcb_4_6]]

[[IMAGE:193adccb815e9fcb_4_7]]

[[IMAGE:193adccb815e9fcb_4_8]]

[[IMAGE:193adccb815e9fcb_4_9]]

A published solution is not available for this question yet.
Question 6 MCQ · 3.0 marks
A student is training a neural network for regression problem with a sigmoid activation at the
output layer [[IMAGE:193adccb815e9fcb_4_10]] to predict the continuous value between 0 and 1. The true mathematical error term
for Mean Squared Error (MSE) loss is
[[IMAGE:193adccb815e9fcb_5_11]] . However, due to a coding mistake, the student implements the error term for Binary Cross-
Entropy (BCE) loss instead, calculating it as [[IMAGE:193adccb815e9fcb_5_12]] . Assuming the network has not yet
reached zero error ( [[IMAGE:193adccb815e9fcb_5_13]] ), which of the following correctly describes the relationship between the
implemented [[IMAGE:193adccb815e9fcb_5_14]] and the true [[IMAGE:193adccb815e9fcb_5_15]] ?






The student's [[IMAGE:193adccb815e9fcb_5_16]] equals the true [[IMAGE:193adccb815e9fcb_5_17]] because the error term [[IMAGE:193adccb815e9fcb_5_18]] always holds
regardless of loss function.



The student's [[IMAGE:193adccb815e9fcb_5_19]] is always at least 4 times larger in magnitude than true [[IMAGE:193adccb815e9fcb_5_20]] .


The student's [[IMAGE:193adccb815e9fcb_5_21]] is smaller than the true [[IMAGE:193adccb815e9fcb_5_22]] because sigmoid squashes the
gradient.


The student's [[IMAGE:193adccb815e9fcb_5_23]] equals the true [[IMAGE:193adccb815e9fcb_5_24]] only when [[IMAGE:193adccb815e9fcb_5_25]] , since [[IMAGE:193adccb815e9fcb_5_26]] is
maximum there and cancels with the MSE factor.




A published solution is not available for this question yet.
Question 7 MCQ · 3.0 marks
An input image of size 32×32 is passed through a CNN containing three convolution layers:
• Conv1: 3x3, stride 1
• Conv2: 3x3, stride 1
• Conv3: 3x3, stride 1
Find the receptive field of a neuron in Conv3.
3x3
4x4
5x5
6x6
7x7
8x8
A published solution is not available for this question yet.
Question 8 MCQ · 3.0 marks
For a convolution layer, suppose the input feature map has size 28x28x64. A 1x1 convolution with
32 filters is applied to it. What will be the output size? (Assume stride = 1 and no padding.)
28x28x64
28x28x32
32x32x28
64x64x32
A published solution is not available for this question yet.
Question 9 MCQ · 3.0 marks
A convolution operation is applied to an input image of size 64x64 using a single 3x3 filter. The
convolution uses:
• Kernel size = 3x3
• Padding = 1
• Stride = 1
What will be the size of the output feature map produced by this filter?
62x62
64x64
66x66
60x60
A published solution is not available for this question yet.
Question 10 MCQ · 3.0 marks
Consider the following PyTorch convolutional layer:
[[IMAGE:193adccb815e9fcb_6_27]]
How many total learnable parameters (including biases) does this layer contain?

4,800
9,600
9,664
4,864
A published solution is not available for this question yet.
Question 11 MCQ · 3.0 marks
Consider the following image preprocessing pipeline:
[[IMAGE:193adccb815e9fcb_7_28]]
Which of the following statements is correct about the above preprocessing pipeline?

Steps A, B, and C add randomness to training images, helping the model
generalize better.
The Normalize step should be placed before converting the image to a tensor
for correct scaling.
Data Augmentation should be added in test pipeline otherwise the code won't
work.
The ColorJitter operation is applied after normalization, so pixel values stay
between 0 and 1.
A published solution is not available for this question yet.
Question 12 MCQ · 2.0 marks
A perceptron in 2D is defined by a linear decision function:
[[IMAGE:193adccb815e9fcb_8_29]]
Suppose a classification rule satisfies:
1. All points with x1 + x2 = 1 are classified as +1
2. All points with x1 + x2 = 0 are classified as -1
3. All points with x1 + x2 = 2 are classified as -1
Which of the following is TRUE?

The problem is linearly separable because classes depend only on x1 + x2
The problem is not linearly separable
A linear separator exists only if w1 = w2
The problem is linearly separable when w1 = -w2
A published solution is not available for this question yet.
Question 13 MCQ · 2.0 marks
Consider the following tensor operation in pytorch:
[[IMAGE:193adccb815e9fcb_8_30]]
What will be the output of the final print statement?

[[IMAGE:193adccb815e9fcb_9_31]]

[[IMAGE:193adccb815e9fcb_9_32]]

[[IMAGE:193adccb815e9fcb_9_33]]

[[IMAGE:193adccb815e9fcb_9_34]]

A published solution is not available for this question yet.
Question 14 MCQ · 2.0 marks
A student is using a single neuron with activation
[[IMAGE:193adccb815e9fcb_9_35]]
where
[[IMAGE:193adccb815e9fcb_9_36]]
Can this single neuron represent XOR exactly for the four binary input pairs?


A suitable choice of parameters can represent XOR exactly.
XOR always requires at least one hidden layer, hence it can not represent XOR
exactly for the four binary input pairs.
It can not represent XOR exactly for the four binary input pairs but only when
[[IMAGE:193adccb815e9fcb_9_37]] .

It can not represent XOR exactly for the four binary input pairs because a
single neuron can only represent linear decision boundaries.
A published solution is not available for this question yet.
Question 15 MCQ · 2.0 marks
Consider the following snippet of a CNN block:
[[IMAGE:193adccb815e9fcb_10_38]]
A part of which classical CNN architecture is represented by the above code?

VGG
GoogleNet
Resnet
Alexnet
A published solution is not available for this question yet.
Question 16 MCQ · 2.0 marks
In deep CNN architectures like VGGNet, multiple 3×3 convolutional layers are stacked instead of
using a single large kernel (e.g.,7×7).
Which of the following statements best explains the advantage of this design choice?
It increases the receptive field and parameter count simultaneously,
improving representational power.
It provides a similar effective receptive field as a larger kernel while
introducing more non-linearities and fewer parameters.
It reduces the receptive field size, but compensates with more skip
connections to retain context.
It limits model depth to avoid overfitting by reducing the total number of
convolutional layers.
A published solution is not available for this question yet.
Question 17 MCQ · 2.0 marks
A CNN produces an output feature map of size 10x10x64. Global Average Pooling (GAP) is applied
to this feature map. What is the dimension of the output after applying the GAP?
10×10×64
1×1×64
1×1×1
64×64×64
A published solution is not available for this question yet.
Question 18 NAT · 3.0 marks
Consider the following neural network and answer the given subquestions.
[[IMAGE:193adccb815e9fcb_11_39]]
A feedforward neural network has 2 input neurons, one hidden layer (l = 1) with 2 neurons, and
one output neuron (l = 2). Every neuron uses the sigmoid activation [[IMAGE:193adccb815e9fcb_11_40]] .
For each layer [[IMAGE:193adccb815e9fcb_11_41]] :
[[IMAGE:193adccb815e9fcb_11_42]]
[[IMAGE:193adccb815e9fcb_11_43]] ,
[[IMAGE:193adccb815e9fcb_11_44]]
Here [[IMAGE:193adccb815e9fcb_11_45]] is the input. Binary cross entropy loss is being used.
[[IMAGE:193adccb815e9fcb_12_46]]
[[IMAGE:193adccb815e9fcb_12_47]]
[[IMAGE:193adccb815e9fcb_12_48]]
Compute the pre-activation input to the output neuron, [[IMAGE:193adccb815e9fcb_12_49]] , and the final output [[IMAGE:193adccb815e9fcb_12_50]] . Give
[[IMAGE:193adccb815e9fcb_12_51]] correct to four decimal places.













A published solution is not available for this question yet.
Question 19 NAT · 4.0 marks
Consider the following neural network and answer the given subquestions.
[[IMAGE:193adccb815e9fcb_11_39]]
A feedforward neural network has 2 input neurons, one hidden layer (l = 1) with 2 neurons, and
one output neuron (l = 2). Every neuron uses the sigmoid activation [[IMAGE:193adccb815e9fcb_11_40]] .
For each layer [[IMAGE:193adccb815e9fcb_11_41]] :
[[IMAGE:193adccb815e9fcb_11_42]]
[[IMAGE:193adccb815e9fcb_11_43]] ,
[[IMAGE:193adccb815e9fcb_11_44]]
Here [[IMAGE:193adccb815e9fcb_11_45]] is the input. Binary cross entropy loss is being used.
[[IMAGE:193adccb815e9fcb_12_46]]
[[IMAGE:193adccb815e9fcb_12_47]]
[[IMAGE:193adccb815e9fcb_12_48]]
[[IMAGE:193adccb815e9fcb_12_52]]
Compute i.e., the gradient of the loss with respect to the hidden-layer weight [[IMAGE:193adccb815e9fcb_12_53]]
(connecting input [[IMAGE:193adccb815e9fcb_12_54]] to hidden neuron [[IMAGE:193adccb815e9fcb_12_55]] ).Give the value correct upto four decimal places.














A published solution is not available for this question yet.
Question 20 MCQ · 1.0 marks
Which statement is guaranteed by the Universal Approximation Theorem (UAT)?
A sufficiently large neural network can always be trained successfully using
gradient descent.
The theorem provides a bound on the number of neurons required for a given
approximation error.
The theorem provides a bound on the amount of training data required.
None of these
A published solution is not available for this question yet.