MauryaHub PYQ Practice

da5006_2026T2_ET_FN.pdf

Deep Learning for Computer Vision · End Term · May 2026 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 2.0 marks

An RGB image of size 128×96 is stored with 32-bit floating-point values for every channel. Ignoring metadata, how many bytes are required?
  1. 36,864
  2. 73,728
  3. 147,456
  4. 294,912

A published solution is not available for this question yet.

Question 3 MCQ · 2.0 marks

Let a 1-D filter be h=[2,-1,3]. A library routine performs cross-correlation. Which filter should be supplied to the routine to reproduce mathematical convolution with h, ignoring boundary effects?
  1. [2,-1,3]
  2. [-2,1,-3]
  3. [3,-1,2]
  4. [3,1,2]

A published solution is not available for this question yet.

Question 4 MCQ · 2.0 marks

At a point on an ideal straight intensity edge, the local image-gradient vector is most naturally interpreted as pointing:
  1. Along the tangent direction of the edge
  2. Approximately normal to the edge
  3. Along the direction of maximum smoothing
  4. Along the local isophote

A published solution is not available for this question yet.

Question 5 MCQ · 2.0 marks

An image is smoothed successively with Gaussian kernels of standard deviations 2 and 3. Ignoring discretization, the equivalent single Gaussian has standard deviation:
  1. [[IMAGE:1441985235580fa1_3_2]]
    Source diagram or notation
  2. [[IMAGE:1441985235580fa1_3_3]]
    Source diagram or notation
  3. [[IMAGE:1441985235580fa1_3_4]]
    Source diagram or notation
  4. [[IMAGE:1441985235580fa1_3_5]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 6 MCQ · 2.0 marks

Which design choice is most directly responsible for the rotation robustness of a SIFT descriptor?
  1. Detecting candidate keypoints across multiple image scales
  2. Expressing local gradient orientations relative to an assigned keypoint orientation
  3. Normalizing the descriptor vector after histogram construction
  4. Building an image pyramid without assigning an orientation

A published solution is not available for this question yet.

Question 7 MCQ · 2.0 marks

A hidden activation h is used by two downstream branches, and both branches affect the same scalar loss L. During backpropagation, the gradient with respect to h is obtained by:
  1. Taking the larger of the two branch gradients
  2. Summing the gradient contributions from the two branches
  3. Multiplying the two branch gradients
  4. Averaging the two branch gradients regardless of the graph

A published solution is not available for this question yet.

Question 8 MCQ · 2.0 marks

Suppose successive gradients keep pointing in a similar direction over several updates. Compared with plain gradient descent using the same nominal learning rate, momentum tends to:
  1. Cancel updates in that direction
  2. Build velocity in that direction while smoothing short-term gradient fluctuations
  3. Make the gradient exactly unbiased
  4. Force every parameter to use the same accumulated gradient

A published solution is not available for this question yet.

Question 9 MCQ · 2.0 marks

Training loss continues to decrease, but validation loss has begun to increase consistently. Which interpretation is most appropriate?
  1. The model is necessarily underfitting
  2. Generalization is worsening even though optimization on the training set is improving
  3. The training objective is necessarily non-differentiable
  4. The model has reached the global optimum on unseen data

A published solution is not available for this question yet.

Question 10 MCQ · 2.0 marks

A 3×3 convolution maps 16 input channels to 32 output channels. If the input spatial size changes from 64×64 to 128×128 while the layer configuration is unchanged, the number of trainable convolution weights:
  1. Doubles
  2. Quadruples
  3. Remains unchanged
  4. Is halved

A published solution is not available for this question yet.

Question 11 MCQ · 2.0 marks

A convolution kernel is reused at many spatial locations. The gradient of the loss with respect to one kernel coefficient therefore contains:
  1. A contribution from only the central output location
  2. Contributions accumulated from all output locations where that coefficient was used
  3. Only the gradient of the bias corresponding to that output channel
  4. A contribution from exactly one spatial location chosen during the forward pass

A published solution is not available for this question yet.

Question 12 MCQ · 2.0 marks

[[IMAGE:1441985235580fa1_5_6]]
Source diagram or notation
  1. It forces F(x) to be constant
  2. The block can approach an identity mapping when the residual F(x) approaches zero
  3. It prevents gradients from entering the residual branch
  4. It makes every convolution invertible

A published solution is not available for this question yet.

Question 13 MCQ · 2.0 marks

A depthwise-separable convolution first applies a spatial filter independently to each input channel and then uses a 1×1 convolution across channels. Its main computational advantage over a dense spatial convolution is that it:
  1. Eliminates channel mixing from the layer
  2. Factorizes spatial filtering and channel mixing into cheaper operations
  3. Uses a larger spatial kernel with the same parameter count
  4. Forces the number of output channels to equal the number of input channels

A published solution is not available for this question yet.

Question 14 MCQ · 2.0 marks

When adapting a pretrained CNN to a related task with a relatively small labeled dataset, which training choice is generally the more conservative starting point?
  1. Reinitialize the entire network and use a very large learning rate
  2. Retain pretrained features and update them cautiously, often with a smaller learning rate than newly initialized layers
  3. Freeze the newly added prediction head and update only the old classifier
  4. Randomly permute pretrained channels before optimization

A published solution is not available for this question yet.

Question 15 MCQ · 2.0 marks

Which processing pattern is characteristic of a two-stage object detector?
  1. Dense prediction of final detections in one stage with no proposal-processing step
  2. Generation of candidate regions followed by prediction on those candidate regions
  3. Pixel-wise class prediction without object localization
  4. Sequence-level prediction using recurrent hidden states

A published solution is not available for this question yet.

Question 16 MCQ · 2.0 marks

For semantic image segmentation with C classes, the model output before the final class decision is naturally organized as:
  1. One C-dimensional vector for the entire image only
  2. A spatial grid with C class scores at each output location
  3. One bounding box for each of the C classes
  4. A sequence with one recurrent state per class and no spatial layout

A published solution is not available for this question yet.

Question 17 MCQ · 2.0 marks

Compared with a basic recurrent update, the gating mechanisms in LSTMs and GRUs are primarily intended to:
  1. Remove nonlinearities from the recurrent computation
  2. Regulate information retention and update across timesteps
  3. Make the model invariant to permutation of the sequence
  4. Replace recurrent state with a fixed convolution kernel

A published solution is not available for this question yet.

Question 18 MCQ · 2.0 marks

In self-attention, changing the query vector for one token while keeping keys and values fixed directly changes:
  1. The number of tokens in the sequence
  2. The attention weights used to combine the value vectors for that token
  3. The dimensionality of every value vector
  4. The number of transformer layers

A published solution is not available for this question yet.

Question 19 MCQ · 2.0 marks

A standard way to convert an image for processing by a Vision Transformer is to:
  1. Treat the complete image batch as a single token
  2. Divide the image into patches and map the patches to token embeddings
  3. Replace each patch with its ground-truth class before attention
  4. Use only the global mean RGB value as the token sequence

A published solution is not available for this question yet.

Question 20 MCQ · 2.0 marks

Which statement correctly distinguishes the two model families?
  1. A GAN necessarily contains an explicit encoder that outputs a Gaussian posterior
  2. A VAE uses a latent-variable objective with a regularized approximate posterior, whereas a GAN uses an adversarial training objective
  3. A VAE is trained only through a discriminator
  4. GANs and VAEs have the same training objective but different optimizers

A published solution is not available for this question yet.

Question 21 MCQ · 2.0 marks

Classifier-free guidance during sampling is based on combining:
  1. Predictions from two external classifiers trained on different label sets
  2. Conditional and unconditional denoising predictions from the diffusion model
  3. The forward-noise sample and a segmentation mask with equal weights
  4. Two independently sampled latent codes without conditioning information

A published solution is not available for this question yet.

Question 22 MSQ · 3.0 marks

Select all correct statements
  1. A normalized averaging filter preserves a constant-valued region away from boundary effects.
  2. Increasing Gaussian scale suppresses progressively finer image structure.
  3. Image pyramids represent image information at multiple spatial scales.
  4. Cross-correlation always flips the filter before applying it.

A published solution is not available for this question yet.

Question 23 MSQ · 3.0 marks

Select all correct statements .
  1. Backpropagation applies the chain rule efficiently through a computation graph.
  2. Regularization can change the learned solution even when the network architecture is unchanged.
  3. Gradient-descent variants can modify how current and past gradient information is used to update parameters.
  4. Improving training means validation performance must increase after every individual parameter update.

A published solution is not available for this question yet.

Question 24 MSQ · 3.0 marks

Select all correct statements.
  1. Weight sharing allows the same convolution kernel parameters to be used at multiple spatial locations.
  2. Inception-style designs can use multiple processing branches.
  3. Residual connections create shortcut paths across learned transformations.
  4. Finetuning requires discarding all pretrained weights before optimization starts.

A published solution is not available for this question yet.

Question 25 MSQ · 3.0 marks

Select all correct statements.
  1. Two-stage detectors separate candidate-region generation from later prediction on candidates.
  2. Single-stage detectors can produce final detection predictions without a separate proposal-processing stage.
  3. Segmentation is a dense prediction task with spatially resolved outputs.
  4. Object detection and segmentation necessarily produce identical output representations.

A published solution is not available for this question yet.

Question 26 MSQ · 3.0 marks

Select all correct statements .
  1. Recurrent models can process ordered sequences of visual features.
  2. Backpropagation through a recurrent model propagates gradient information across unrolled timesteps.
  3. LSTMs and GRUs use gating mechanisms.
  4. Randomly permuting all video frames leaves every temporal model output unchanged by definition.

A published solution is not available for this question yet.

Question 27 MSQ · 3.0 marks

Select all correct statements.
  1. Soft attention can form a weighted combination of visual features.
  2. Self-attention allows token representations to depend directly on other tokens.
  3. Vision Transformers adapt transformer-style token processing to visual inputs.
  4. Transformer-based segmentation is restricted to producing one class label for an entire image.

A published solution is not available for this question yet.

Question 28 MSQ · 3.0 marks

Select all correct statements.
  1. GAN training involves adversarial objectives.
  2. VAEs introduce latent variables and regularize an approximate posterior.
  3. DDPMs use a noise-based forward process together with a learned reverse denoising process.
  4. Classifier-free guidance requires a separate external classifier at sampling time.

A published solution is not available for this question yet.

Question 29 MSQ · 3.0 marks

Select all correct statements .
  1. SimCLR is a self-supervised contrastive representation-learning method.
  2. MoCo is a contrastive representation-learning method.
  3. CLIP learns relationships between visual and textual representations.
  4. In SimCLR, a positive pair is formed by two unrelated images solely because they receive the same predicted class.

A published solution is not available for this question yet.

Question 30 NAT · 4.0 marks

An input feature map has spatial size 63×63. A 5×5 convolution is applied with stride 2 and padding 1. Compute the output size along one spatial dimension.

    A published solution is not available for this question yet.

    Question 31 NAT · 4.0 marks

    An image is smoothed successively by Gaussian kernels with σ1=1.2 and σ2=1.6. Ignoring discretization, compute the equivalent single Gaussian standard deviation.

      A published solution is not available for this question yet.

      Question 32 NAT · 4.0 marks

      [[IMAGE:1441985235580fa1_11_7]]
      Source diagram or notation

        A published solution is not available for this question yet.

        Question 33 NAT · 4.0 marks

        A convolutional layer uses 3×3 kernels, 16 input channels and 32 output channels. The layer has one bias per output channel. Compute the total number of trainable parameters.

          A published solution is not available for this question yet.

          Question 34 NAT · 4.0 marks

          A depthwise-separable convolution receives 32 channels. It uses a 3×3 depthwise convolution followed by a 1×1 pointwise convolution producing 64 output channels. Ignore biases. Compute the total number of weights.

            A published solution is not available for this question yet.

            Question 35 NAT · 4.0 marks

            [[IMAGE:1441985235580fa1_12_8]]
            Source diagram or notation

              A published solution is not available for this question yet.

              Question 36 NAT · 4.0 marks

              [[IMAGE:1441985235580fa1_12_9]]
              Source diagram or notation

                A published solution is not available for this question yet.

                Question 37 NAT · 4.0 marks

                [[IMAGE:1441985235580fa1_13_10]]
                Source diagram or notation

                  A published solution is not available for this question yet.

                  Question 38 NAT · 4.0 marks

                  [[IMAGE:1441985235580fa1_13_11]]
                  Source diagram or notation

                    A published solution is not available for this question yet.