bt4001_2026T1_ET_FN.pdf
Algorithmic Thinking in Bioinformatics · End Term · Jan 2026 FN
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MCQ · 5.0 marks
You are given two DNA sequences [[IMAGE:a5f980a496e7976b_2_2]] and [[IMAGE:a5f980a496e7976b_2_3]] of length [[IMAGE:a5f980a496e7976b_2_4]] and [[IMAGE:a5f980a496e7976b_2_5]]
respectively. Assume [[IMAGE:a5f980a496e7976b_2_6]] to be the scoring function. A pseudocode for an algorithm that
calculates the optimal global alignment score between the two sequences is given below. Fill in
the blanks for the algorithm appropriately.
[[IMAGE:a5f980a496e7976b_2_7]]






[[IMAGE:a5f980a496e7976b_2_8]]

[[IMAGE:a5f980a496e7976b_2_9]]

[[IMAGE:a5f980a496e7976b_2_10]]

[[IMAGE:a5f980a496e7976b_3_11]]

A published solution is not available for this question yet.
Question 3 MCQ · 5.0 marks
The number of occurrences of the four nucleotides in the genome sequence of some virus is given
in Table 1.
[[IMAGE:a5f980a496e7976b_3_12]]
The suffix array of the genome sequence is: [[IMAGE:a5f980a496e7976b_3_13]] . Assume $ is present at
the end of the original genome sequence for all computations. Which of the following represents
the correct genome sequence?


AAGGGCACTG$
GTGCCGAGAA$
CGATCGGAGA$
GTGCGCAGAA$
The genome sequence cannot be derived with the information provided.
A published solution is not available for this question yet.
Question 4 MCQ · 5.0 marks
Consider the following profile matrix.
[[IMAGE:a5f980a496e7976b_3_14]]
Find the profile-most probable [[IMAGE:a5f980a496e7976b_3_15]] -mer for the sequence [[IMAGE:a5f980a496e7976b_3_16]] .



GCTA
GGCT
CTAT
TATA
A published solution is not available for this question yet.
Question 5 MCQ · 5.0 marks
Let [[IMAGE:a5f980a496e7976b_4_17]] be a random sample from a distribution with density:
[[IMAGE:a5f980a496e7976b_4_18]]
Find the maximum likelihood estimator (MLE) of [[IMAGE:a5f980a496e7976b_4_19]] .



[[IMAGE:a5f980a496e7976b_4_20]]

[[IMAGE:a5f980a496e7976b_4_21]]

[[IMAGE:a5f980a496e7976b_4_22]]

[[IMAGE:a5f980a496e7976b_4_23]]

A published solution is not available for this question yet.
Question 6 MSQ · 5.0 marks
A mass spectrometer generates the following spectrum [[IMAGE:a5f980a496e7976b_4_24]] for the peptide '' [[IMAGE:a5f980a496e7976b_4_25]]
HEQN''.
[[IMAGE:a5f980a496e7976b_4_26]]
The figure given below depicts the mass of all amino acids.
[[IMAGE:a5f980a496e7976b_5_27]]
The criteria to trim peptides in the leaderboard based cyclopeptide sequencing algorithm is
modified as follows:
"Suppose the algorithm is considering the peptide [[IMAGE:a5f980a496e7976b_5_28]] . Let the number of elements in the
theoretical spectrum of [[IMAGE:a5f980a496e7976b_5_29]] be [[IMAGE:a5f980a496e7976b_5_30]] . If [[IMAGE:a5f980a496e7976b_5_31]] ( [[IMAGE:a5f980a496e7976b_5_32]] , [[IMAGE:a5f980a496e7976b_5_33]] ) [[IMAGE:a5f980a496e7976b_5_34]] , then the peptide is
trimmed''.
Which of the following peptides will be trimmed?











ENQ
EQT
HE
NE
None of these
A published solution is not available for this question yet.
Question 7 NAT · 5.0 marks
Given below are a set of genetic sequences and the position (sites) of their nucleotides.
[[IMAGE:a5f980a496e7976b_5_35]]
You construct the following rooted binary phylogenetic tree from this. Calculate the parsimony
score for the tree.
Enter the value as a single integer.
[[IMAGE:a5f980a496e7976b_6_36]]


A published solution is not available for this question yet.
Question 8 NAT · 5.0 marks
Consider the below table containing sample data from a sequencing experiment. Each trial
corresponds to 5 independent expression measurements. Each measurement is classified as
either:
• H: high expression signal
• L: low expression signal
Assume that each trial is generated from one of two hidden biological states, each with its own
probability of producing a high signal.
[[IMAGE:a5f980a496e7976b_6_37]]
Let [[IMAGE:a5f980a496e7976b_6_38]] and [[IMAGE:a5f980a496e7976b_6_39]] denote the probability of observing a high signal (H) in biological states 1 and 2,
respectively.
Given:
[[IMAGE:a5f980a496e7976b_6_40]]
Here, π represents the prior probability of the two biological states.
What is the probability that **Trial 3** was generated by biological state 1?
Answer correct to two decimal places.




A published solution is not available for this question yet.
Question 9 NAT · 4.0 marks
A researcher is analyzing gene expression levels from five tissue samples. Gene expression
measures how active a gene is in a given sample; higher values mean the gene is more active. The
researcher believes the samples come from three different biological subtypes and uses a
univariate Gaussian Mixture Model (GMM) with 3 components to model them.
Suppose for component 2, you are given the following expression levels and corresponding
responsibilities. Calculate the value of [[IMAGE:a5f980a496e7976b_7_41]] ?
[[IMAGE:a5f980a496e7976b_7_42]]
Answer correct to two decimal places.


A published solution is not available for this question yet.
Question 10 NAT · 2.0 marks
The overlap graph of some linear (i.e., non-circular) genome [[IMAGE:a5f980a496e7976b_8_43]] has [[IMAGE:a5f980a496e7976b_8_44]] nodes. Answer the
given subquestions about the de Bruijn graph constructed on the same set of k-mers. Remember
that [[IMAGE:a5f980a496e7976b_8_45]] can be reconstructed by using both the graphs when answering these questions.
What is the number of edges in the de Bruijn graph?
Enter the value as a single integer.



A published solution is not available for this question yet.
Question 11 NAT · 3.0 marks
The overlap graph of some linear (i.e., non-circular) genome [[IMAGE:a5f980a496e7976b_8_43]] has [[IMAGE:a5f980a496e7976b_8_44]] nodes. Answer the
given subquestions about the de Bruijn graph constructed on the same set of k-mers. Remember
that [[IMAGE:a5f980a496e7976b_8_45]] can be reconstructed by using both the graphs when answering these questions.
What is the maximum number of nodes possible in the de Bruijn graph?
Enter the value as a single integer.



A published solution is not available for this question yet.
Question 12 NAT · 3.0 marks
Consider a dataset of six genes. Each gene is represented by its expression levels under two
experimental conditions (Condition A and Condition B), so each gene corresponds to a point in [[IMAGE:a5f980a496e7976b_9_46]] .
[[IMAGE:a5f980a496e7976b_9_47]]
A researcher applies **K-means clustering** with [[IMAGE:a5f980a496e7976b_9_48]] to group genes with similar expression
patterns.
The cluster centroids are initialized as follows:
• Cluster 1 centroid: [[IMAGE:a5f980a496e7976b_9_49]] (low-expression genes)
• Cluster 2 centroid: [[IMAGE:a5f980a496e7976b_9_50]] (high-expression genes)
Based on the above data, answer the given subquestions.
If [[IMAGE:a5f980a496e7976b_9_51]] is the centroid of the first cluster after convergence, find the value of [[IMAGE:a5f980a496e7976b_9_52]] .
Enter the value as a single integer.







A published solution is not available for this question yet.
Question 13 NAT · 3.0 marks
Consider a dataset of six genes. Each gene is represented by its expression levels under two
experimental conditions (Condition A and Condition B), so each gene corresponds to a point in [[IMAGE:a5f980a496e7976b_9_46]] .
[[IMAGE:a5f980a496e7976b_9_47]]
A researcher applies **K-means clustering** with [[IMAGE:a5f980a496e7976b_9_48]] to group genes with similar expression
patterns.
The cluster centroids are initialized as follows:
• Cluster 1 centroid: [[IMAGE:a5f980a496e7976b_9_49]] (low-expression genes)
• Cluster 2 centroid: [[IMAGE:a5f980a496e7976b_9_50]] (high-expression genes)
Based on the above data, answer the given subquestions.
If [[IMAGE:a5f980a496e7976b_10_53]] is the centroid of the second cluster after convergence, find the value of [[IMAGE:a5f980a496e7976b_10_54]] .
Enter the value as a single integer.







A published solution is not available for this question yet.