MauryaHub PYQ Practice

bt4001_2026T1_ET_FN.pdf

Algorithmic Thinking in Bioinformatics · End Term · Jan 2026 FN

← Course papers · Start practice / exam

Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.

Question 2 MCQ · 5.0 marks

You are given two DNA sequences [[IMAGE:a5f980a496e7976b_2_2]] and [[IMAGE:a5f980a496e7976b_2_3]] of length [[IMAGE:a5f980a496e7976b_2_4]] and [[IMAGE:a5f980a496e7976b_2_5]] respectively. Assume [[IMAGE:a5f980a496e7976b_2_6]] to be the scoring function. A pseudocode for an algorithm that calculates the optimal global alignment score between the two sequences is given below. Fill in the blanks for the algorithm appropriately. [[IMAGE:a5f980a496e7976b_2_7]]
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. [[IMAGE:a5f980a496e7976b_2_8]]
    Source diagram or notation
  2. [[IMAGE:a5f980a496e7976b_2_9]]
    Source diagram or notation
  3. [[IMAGE:a5f980a496e7976b_2_10]]
    Source diagram or notation
  4. [[IMAGE:a5f980a496e7976b_3_11]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 3 MCQ · 5.0 marks

The number of occurrences of the four nucleotides in the genome sequence of some virus is given in Table 1. [[IMAGE:a5f980a496e7976b_3_12]] The suffix array of the genome sequence is: [[IMAGE:a5f980a496e7976b_3_13]] . Assume $ is present at the end of the original genome sequence for all computations. Which of the following represents the correct genome sequence?
Source diagram or notationSource diagram or notation
  1. AAGGGCACTG$
  2. GTGCCGAGAA$
  3. CGATCGGAGA$
  4. GTGCGCAGAA$
  5. The genome sequence cannot be derived with the information provided.

A published solution is not available for this question yet.

Question 4 MCQ · 5.0 marks

Consider the following profile matrix. [[IMAGE:a5f980a496e7976b_3_14]] Find the profile-most probable [[IMAGE:a5f980a496e7976b_3_15]] -mer for the sequence [[IMAGE:a5f980a496e7976b_3_16]] .
Source diagram or notationSource diagram or notationSource diagram or notation
  1. GCTA
  2. GGCT
  3. CTAT
  4. TATA

A published solution is not available for this question yet.

Question 5 MCQ · 5.0 marks

Let [[IMAGE:a5f980a496e7976b_4_17]] be a random sample from a distribution with density: [[IMAGE:a5f980a496e7976b_4_18]] Find the maximum likelihood estimator (MLE) of [[IMAGE:a5f980a496e7976b_4_19]] .
Source diagram or notationSource diagram or notationSource diagram or notation
  1. [[IMAGE:a5f980a496e7976b_4_20]]
    Source diagram or notation
  2. [[IMAGE:a5f980a496e7976b_4_21]]
    Source diagram or notation
  3. [[IMAGE:a5f980a496e7976b_4_22]]
    Source diagram or notation
  4. [[IMAGE:a5f980a496e7976b_4_23]]
    Source diagram or notation

A published solution is not available for this question yet.

Question 6 MSQ · 5.0 marks

A mass spectrometer generates the following spectrum [[IMAGE:a5f980a496e7976b_4_24]] for the peptide '' [[IMAGE:a5f980a496e7976b_4_25]] HEQN''. [[IMAGE:a5f980a496e7976b_4_26]] The figure given below depicts the mass of all amino acids. [[IMAGE:a5f980a496e7976b_5_27]] The criteria to trim peptides in the leaderboard based cyclopeptide sequencing algorithm is modified as follows: "Suppose the algorithm is considering the peptide [[IMAGE:a5f980a496e7976b_5_28]] . Let the number of elements in the theoretical spectrum of [[IMAGE:a5f980a496e7976b_5_29]] be [[IMAGE:a5f980a496e7976b_5_30]] . If [[IMAGE:a5f980a496e7976b_5_31]] ( [[IMAGE:a5f980a496e7976b_5_32]] , [[IMAGE:a5f980a496e7976b_5_33]] ) [[IMAGE:a5f980a496e7976b_5_34]] , then the peptide is trimmed''. Which of the following peptides will be trimmed?
Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation
  1. ENQ
  2. EQT
  3. HE
  4. NE
  5. None of these

A published solution is not available for this question yet.

Question 7 NAT · 5.0 marks

Given below are a set of genetic sequences and the position (sites) of their nucleotides. [[IMAGE:a5f980a496e7976b_5_35]] You construct the following rooted binary phylogenetic tree from this. Calculate the parsimony score for the tree. Enter the value as a single integer. [[IMAGE:a5f980a496e7976b_6_36]]
Source diagram or notationSource diagram or notation

    A published solution is not available for this question yet.

    Question 8 NAT · 5.0 marks

    Consider the below table containing sample data from a sequencing experiment. Each trial corresponds to 5 independent expression measurements. Each measurement is classified as either: • H: high expression signal • L: low expression signal Assume that each trial is generated from one of two hidden biological states, each with its own probability of producing a high signal. [[IMAGE:a5f980a496e7976b_6_37]] Let [[IMAGE:a5f980a496e7976b_6_38]] and [[IMAGE:a5f980a496e7976b_6_39]] denote the probability of observing a high signal (H) in biological states 1 and 2, respectively. Given: [[IMAGE:a5f980a496e7976b_6_40]] Here, π represents the prior probability of the two biological states. What is the probability that **Trial 3** was generated by biological state 1? Answer correct to two decimal places.
    Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

      A published solution is not available for this question yet.

      Question 9 NAT · 4.0 marks

      A researcher is analyzing gene expression levels from five tissue samples. Gene expression measures how active a gene is in a given sample; higher values mean the gene is more active. The researcher believes the samples come from three different biological subtypes and uses a univariate Gaussian Mixture Model (GMM) with 3 components to model them. Suppose for component 2, you are given the following expression levels and corresponding responsibilities. Calculate the value of [[IMAGE:a5f980a496e7976b_7_41]] ? [[IMAGE:a5f980a496e7976b_7_42]] Answer correct to two decimal places.
      Source diagram or notationSource diagram or notation

        A published solution is not available for this question yet.

        Question 10 NAT · 2.0 marks

        The overlap graph of some linear (i.e., non-circular) genome [[IMAGE:a5f980a496e7976b_8_43]] has [[IMAGE:a5f980a496e7976b_8_44]] nodes. Answer the given subquestions about the de Bruijn graph constructed on the same set of k-mers. Remember that [[IMAGE:a5f980a496e7976b_8_45]] can be reconstructed by using both the graphs when answering these questions.
        What is the number of edges in the de Bruijn graph? Enter the value as a single integer.
        Source diagram or notationSource diagram or notationSource diagram or notation

          A published solution is not available for this question yet.

          Question 11 NAT · 3.0 marks

          The overlap graph of some linear (i.e., non-circular) genome [[IMAGE:a5f980a496e7976b_8_43]] has [[IMAGE:a5f980a496e7976b_8_44]] nodes. Answer the given subquestions about the de Bruijn graph constructed on the same set of k-mers. Remember that [[IMAGE:a5f980a496e7976b_8_45]] can be reconstructed by using both the graphs when answering these questions.
          What is the maximum number of nodes possible in the de Bruijn graph? Enter the value as a single integer.
          Source diagram or notationSource diagram or notationSource diagram or notation

            A published solution is not available for this question yet.

            Question 12 NAT · 3.0 marks

            Consider a dataset of six genes. Each gene is represented by its expression levels under two experimental conditions (Condition A and Condition B), so each gene corresponds to a point in [[IMAGE:a5f980a496e7976b_9_46]] . [[IMAGE:a5f980a496e7976b_9_47]] A researcher applies **K-means clustering** with [[IMAGE:a5f980a496e7976b_9_48]] to group genes with similar expression patterns. The cluster centroids are initialized as follows: • Cluster 1 centroid: [[IMAGE:a5f980a496e7976b_9_49]] (low-expression genes) • Cluster 2 centroid: [[IMAGE:a5f980a496e7976b_9_50]] (high-expression genes) Based on the above data, answer the given subquestions.
            If [[IMAGE:a5f980a496e7976b_9_51]] is the centroid of the first cluster after convergence, find the value of [[IMAGE:a5f980a496e7976b_9_52]] . Enter the value as a single integer.
            Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

              A published solution is not available for this question yet.

              Question 13 NAT · 3.0 marks

              Consider a dataset of six genes. Each gene is represented by its expression levels under two experimental conditions (Condition A and Condition B), so each gene corresponds to a point in [[IMAGE:a5f980a496e7976b_9_46]] . [[IMAGE:a5f980a496e7976b_9_47]] A researcher applies **K-means clustering** with [[IMAGE:a5f980a496e7976b_9_48]] to group genes with similar expression patterns. The cluster centroids are initialized as follows: • Cluster 1 centroid: [[IMAGE:a5f980a496e7976b_9_49]] (low-expression genes) • Cluster 2 centroid: [[IMAGE:a5f980a496e7976b_9_50]] (high-expression genes) Based on the above data, answer the given subquestions.
              If [[IMAGE:a5f980a496e7976b_10_53]] is the centroid of the second cluster after convergence, find the value of [[IMAGE:a5f980a496e7976b_10_54]] . Enter the value as a single integer.
              Source diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notationSource diagram or notation

                A published solution is not available for this question yet.