Quiz1
Statistics for Data Science I · Quiz 1 · May 2026
← Course papers · Start practice / exam
Questions and published explanations below are available without starting a test. Some questions may not have a published solution yet.
Question 2 MSQ · 2.0 marks
A researcher selects 200 scholars from a research institute and records their registration numbers,
the time spent on their research projects per week (in hours), their monthly stipend (in Rs.),
alongside their preferred mode of conducting research (Laboratory-based, Computer-based, or
Hybrid).
Based on this study, choose all the correct statements from the following:
"All scholars in the institute" represents the population, while "the 200
selected scholars" represents the sample.
The "preferred mode of research" is a qualitative variable, whereas the "time
spent on research per week" is a quantitative variable.
The time spent on research per week is a categorical variable.
The Registration numbers are numerical variables.
Monthly stipend measured in rupees is a numerical variable.
Published solution
**1. Understand / Given:** The researcher studies 200 scholars from a research institute. Four things are recorded: registration number, weekly research time (hours), monthly stipend (Rs.), and preferred mode of research (Laboratory-based, Computer-based, Hybrid).
**2. Apply the concepts:**
- Population = the whole group of interest. Sample = the part of it that is actually studied.
- Qualitative (categorical) variable = takes labels or categories.
- Quantitative (numerical) variable = takes numbers on which arithmetic is meaningful (hours, rupees).
- A registration number is only an identifier (a label), even though it is written with digits.
**3. Check each statement:**
- All scholars in the institute are the population; the 200 selected scholars are the sample. True.
- Preferred mode of research is a category, so it is qualitative. Time spent per week is measured in hours, so it is quantitative. True.
- Time spent per week is a measured number, not a category. So calling it categorical is False.
- Registration numbers are only labels; adding or averaging them has no meaning. So this is False.
- Monthly stipend in rupees is a measured number, so it is numerical. True.
**4. Conclude:** The correct statements are the population/sample statement, the qualitative/quantitative statement, and the stipend statement.
Answer: A, B, E — population/sample statement, qualitative vs quantitative statement, and monthly stipend as a numerical variable.
A: **Correct:** All scholars of the institute form the population, and the 200 selected scholars form the sample.
B: **Correct:** Preferred mode is a category (qualitative); weekly research time in hours is a number (quantitative).
C: **Incorrect:** Time in hours is a measured numerical value, so it is quantitative, not categorical.
D: **Incorrect:** Registration numbers are only identifiers or labels; arithmetic on them has no meaning, so they are not numerical variables.
E: **Correct:** Stipend in rupees is a measured amount on which arithmetic is meaningful, so it is a numerical variable.
Question 3 MSQ · 2.0 marks
Which of the following statements is/are meaningful?
A person ranked 2nd performed twice as well as a person ranked 4th.
A temperature of [[IMAGE:34fe2c26a53013f3_2_2]] is twice as hot as a temperature of [[IMAGE:34fe2c26a53013f3_2_3]] .


A package weighing [[IMAGE:34fe2c26a53013f3_2_4]] kg is twice as heavy as a package weighing [[IMAGE:34fe2c26a53013f3_2_5]] kg.


Blood groups do not have a natural ordering.
Published solution
**1. Concept:** A statement is meaningful only if the comparison it makes is valid for that type of data.
- Ordinal data (ranks) only tells order, not how much better.
- Interval scale (like \(^\circ C\)) has no true zero, so ratios such as "twice as hot" are not meaningful.
- Ratio scale (like weight in kg) has a true zero, so ratios are meaningful.
- Nominal data (like blood groups) has no natural order.
**2. Check each statement:**
- Rank 2 vs rank 4: ranks only give order. Performing "twice as well" is not meaningful.
- \(20^\circ C\) vs \(10^\circ C\): Celsius has an arbitrary zero. In kelvin these are \(293.15\) K and \(283.15\) K, which are not in the ratio \(2:1\). Not meaningful.
- \(10\) kg vs \(5\) kg: weight has a true zero and \(10 = 2\times 5\). Meaningful.
- Blood groups (A, B, AB, O) are nominal categories and have no natural order. Meaningful (true).
**3. Conclude:** The meaningful statements are the package-weight statement and the blood-group statement.
Answer: C, D — the weight comparison and the blood-group statement are meaningful.
A: **Incorrect:** Ranks are ordinal; they show order only, so "twice as well" has no meaning.
B: **Incorrect:** Celsius is an interval scale with no true zero, so \(20^\circ C\) is not "twice as hot" as \(10^\circ C\).
C: **Correct:** Weight is on a ratio scale with a true zero, and \(10\) kg is exactly twice \(5\) kg.
D: **Correct:** Blood groups are nominal categories, so they have no natural ordering.
Question 4 MSQ · 2.0 marks
Choose all the correct statements from the following:
Monthly electricity consumption of a household from January 2025 to
December 2025 is a Time-series data.
Marks obtained by all students in a Statistics exam in Quiz 1 is a Time-series
data.
Monthly salaries of 500 employees in a company recorded in June 2026 is a
cross-sectional data.
Marks obtained by a particular student in Statistics tests conducted every
month during the year is cross sectional data.
Published solution
**1. Concept:**
- Time-series data = observations of the same unit (person, household, company) recorded over different time points.
- Cross-sectional data = observations of many different units recorded at the same point in time.
**2. Check each statement:**
- Monthly electricity use of one household from Jan to Dec 2025: one unit, many time points. This is time-series. True.
- Marks of all students in one exam (Quiz 1): many units, one time point. This is cross-sectional, not time-series. False.
- Monthly salaries of 500 employees recorded in June 2026: many units, one time point. This is cross-sectional. True.
- Marks of one student in tests held every month: one unit, many time points. This is time-series, not cross-sectional. False.
**3. Conclude:** Only the first and third statements are correct.
Answer: A, C — the household electricity data is time-series, and the salaries of 500 employees in June 2026 are cross-sectional.
A: **Correct:** One household observed month after month gives time-series data.
B: **Incorrect:** Marks of many students in a single exam are recorded at one time point, so the data is cross-sectional.
C: **Correct:** Many employees observed at one time (June 2026) give cross-sectional data.
D: **Incorrect:** One student observed over many months gives time-series data, not cross-sectional.
Question 5 NAT · 2.0 marks
[[IMAGE:34fe2c26a53013f3_3_6]] , [[IMAGE:34fe2c26a53013f3_3_7]] , [[IMAGE:34fe2c26a53013f3_3_8]] , and [[IMAGE:34fe2c26a53013f3_3_9]] are observations, and the sum of their frequencies is 200. The relative frequencies
corresponding to [[IMAGE:34fe2c26a53013f3_3_10]] , [[IMAGE:34fe2c26a53013f3_3_11]] , and [[IMAGE:34fe2c26a53013f3_3_12]] are 20%, 25%, and 30%, respectively. Based on this information,
answer the given subquestions:
What is the frequency corresponding to [[IMAGE:34fe2c26a53013f3_3_13]] ?








Published solution
**1. Given:** The sum of the frequencies of \(a, b, c, d\) is \(200\). Relative frequencies of \(a, b, c\) are \(20\%, 25\%, 30\%\).
**2. Concept:** The relative frequencies of all observations add up to \(100\%\). Frequency \(=\) relative frequency \(\times\) total.
**3. Find the relative frequency of \(d\):**
\[
100\% - (20\% + 25\% + 30\%) = 100\% - 75\% = 25\%
\]
**4. Find the frequency of \(d\):**
\[
f_d = 0.25 \times 200 = 50
\]
Answer: \(50\).
Question 6 NAT · 3.0 marks
[[IMAGE:34fe2c26a53013f3_3_6]] , [[IMAGE:34fe2c26a53013f3_3_7]] , [[IMAGE:34fe2c26a53013f3_3_8]] , and [[IMAGE:34fe2c26a53013f3_3_9]] are observations, and the sum of their frequencies is 200. The relative frequencies
corresponding to [[IMAGE:34fe2c26a53013f3_3_10]] , [[IMAGE:34fe2c26a53013f3_3_11]] , and [[IMAGE:34fe2c26a53013f3_3_12]] are 20%, 25%, and 30%, respectively. Based on this information,
answer the given subquestions:
What will be the cumulative frequency of [[IMAGE:34fe2c26a53013f3_4_14]] , [[IMAGE:34fe2c26a53013f3_4_15]] , and [[IMAGE:34fe2c26a53013f3_4_16]] ?










A published solution is not available for this question yet.
Question 7 MCQ · 1.0 marks
Figure 1 represents the placement percentage of students in different sectors from an
[[IMAGE:34fe2c26a53013f3_4_17]]
engineering college.
Based on the given data, answer the given subquestions:
What is the mode of the placement sectors?

Analytics
Software
Consultancy
Core
Published solution
**1. Concept:** The mode is the category with the highest frequency (the largest percentage in the pie chart).
**2. Read the chart:** Analytics \(=20\%\), Software \(=35\%\), Consultancy \(=20\%\), Core \(=25\%\).
**3. Compare:** The largest value is \(35\%\), which belongs to Software.
Answer: B — Software.
A: **Incorrect:** Analytics is only \(20\%\), which is not the highest.
B: **Correct:** Software has the highest share, \(35\%\), so it is the mode.
C: **Incorrect:** Consultancy is only \(20\%\), which is not the highest.
D: **Incorrect:** Core is \(25\%\), which is less than \(35\%\).
Question 8 MCQ · 3.0 marks
Figure 1 represents the placement percentage of students in different sectors from an
[[IMAGE:34fe2c26a53013f3_4_17]]
engineering college.
Based on the given data, answer the given subquestions:
If 1000 students were placed according to the percentages shown in the pie chart, which of the
following statements is true?

The difference between the number of students placed in Core and the sum of
those in Consultancy and Analytics is 100.
The mode of the placement sectors is shared by Analytics and Software.
The sector with the second-highest placement has 50 more students placed
than the lowest sector.
The sector with the least placement has half as many students as the highest
placed sector.
Published solution
**1. Given:** \(1000\) students are placed as per the pie chart: Analytics \(20\%\), Software \(35\%\), Consultancy \(20\%\), Core \(25\%\).
**2. Convert percentages to numbers (multiply by \(1000\)):**
\[
\text{Analytics}=200,\ \text{Software}=350,\ \text{Consultancy}=200,\ \text{Core}=250
\]
**3. Check each statement:**
- Core \(-\) (Consultancy \(+\) Analytics) \(=250-(200+200)=-150\). The difference is \(150\), not \(100\). False.
- The mode is Software only (\(350\) is the single highest). It is not shared with Analytics. False.
- Second-highest is Core \(=250\); lowest is \(200\) (Analytics or Consultancy). \(250-200=50\). True.
- Lowest \(=200\); half of highest \(=350/2=175\). \(200\ne175\). False.
**4. Conclude:** Only the third statement is true.
Answer: C — The sector with the second-highest placement has 50 more students than the lowest sector.
A: **Incorrect:** The difference is \(|250-400|=150\), not \(100\).
B: **Incorrect:** Only Software (\(350\)) is the mode; Analytics has only \(200\).
C: **Correct:** Second-highest (Core) is \(250\) and the lowest is \(200\), a difference of \(50\).
D: **Incorrect:** Half of the highest is \(175\), but the lowest is \(200\).
Question 9 NAT · 2.0 marks
600 students are classified according to their level of attendance and academic performance. The
results are given in Table 1.
[[IMAGE:34fe2c26a53013f3_5_18]]
Based on the above data, answer the given subquestions.
What proportion of total students are irregular? (Enter the answer correct to 1 decimal places)

Published solution
**1. Given (Table 1):**
- Regular: \(180 + 120 + 60 = 360\)
- Irregular: \(50 + 90 + 100 = 240\)
- Total students \(=600\)
**2. Formula:** Proportion \(=\dfrac{\text{number of irregular students}}{\text{total students}}\).
**3. Substitute:**
\[
\frac{240}{600} = 0.4
\]
Answer: \(0.4\).
Question 10 NAT · 2.0 marks
600 students are classified according to their level of attendance and academic performance. The
results are given in Table 1.
[[IMAGE:34fe2c26a53013f3_5_18]]
Based on the above data, answer the given subquestions.
What proportion of irregular students are having low academic performance? (Enter the answer
correct to 2 decimal places)

Published solution
**1. Given (Table 1):** Irregular students \(=50+90+100=240\). Irregular students with low performance \(=100\).
**2. Formula:** Proportion \(=\dfrac{\text{irregular and low performance}}{\text{all irregular students}}\). Here the base is the irregular students only (not all \(600\)).
**3. Substitute and calculate:**
\[
\frac{100}{240} = 0.4167\ldots \approx 0.42
\]
Answer: \(0.42\).
Question 11 MSQ · 4.0 marks
In a deck, there are cards numbered [[IMAGE:34fe2c26a53013f3_6_19]] to [[IMAGE:34fe2c26a53013f3_6_20]] such that the number of cards of a particular number
in the deck is same as the number on the card. Which of the following statement(s) is/are true
about the mean and mode of the numbers on this deck of card?
Hint:
[[IMAGE:34fe2c26a53013f3_6_21]]



Mode is not defined for this data.
Mean is [[IMAGE:34fe2c26a53013f3_7_22]] .

Mode is [[IMAGE:34fe2c26a53013f3_7_23]] .

Mode is [[IMAGE:34fe2c26a53013f3_7_24]] .

Mean is [[IMAGE:34fe2c26a53013f3_7_25]] .

Mean is [[IMAGE:34fe2c26a53013f3_7_26]] .

Published solution
**1. Given:** Cards numbered \(1\) to \(44\). The number \(k\) appears \(k\) times in the deck.
**2. Mode:** The mode is the most frequent number. Number \(44\) appears \(44\) times, more than any other number. So the mode is \(44\).
**3. Mean:** Mean \(=\dfrac{\text{sum of all card values}}{\text{total number of cards}}\).
Total number of cards:
\[
\sum_{k=1}^{44} k = \frac{44\times45}{2} = 990
\]
Sum of all card values (each \(k\) appears \(k\) times, so we add \(k\cdot k\)):
\[
\sum_{k=1}^{44} k^2 = \frac{44\times45\times89}{6} = 29370
\]
**4. Calculate the mean:**
\[
\text{Mean}=\frac{29370}{990}=29.666\ldots\approx 29.67
\]
**5. Conclude:** Mode \(=44\) and mean \(\approx 29.67\). (The value \(22.5\) is the plain average of \(1,\ldots,44\), which ignores how many times each number appears.)
Answer: C, F — Mode is \(44\), and Mean is \(29.67\).
A: **Incorrect:** The mode exists: \(44\) appears most often (\(44\) times).
B: **Incorrect:** \(22.5\) is the unweighted average of \(1\) to \(44\); the deck has repeated cards, so the mean is \(29.67\).
C: **Correct:** Number \(44\) occurs \(44\) times, the highest frequency.
D: **Incorrect:** Number \(43\) occurs \(43\) times, fewer than \(44\).
E: **Incorrect:** The mean is \(29370/990\approx29.67\), not \(44\).
F: **Correct:** Mean \(=29370/990\approx29.67\).
Question 12 MSQ · 4.0 marks
A total of [[IMAGE:34fe2c26a53013f3_7_27]] students wrote a competitive exam. Rahul's score is denoted by [[IMAGE:34fe2c26a53013f3_7_28]] , which happens to
be exactly at the [[IMAGE:34fe2c26a53013f3_7_29]] percentile. Suppose the examiner realizes that there was a grading error on
Rahul's paper, and his score [[IMAGE:34fe2c26a53013f3_7_30]] needs to increase slightly but remains below [[IMAGE:34fe2c26a53013f3_7_31]] percentile values;
Choose all the correct statements from the following:





In the original results, approximately [[IMAGE:34fe2c26a53013f3_7_32]] students scored less than or equal
to [[IMAGE:34fe2c26a53013f3_7_33]] .


Changing the value of [[IMAGE:34fe2c26a53013f3_7_34]] will definitely alter the mean score of the [[IMAGE:34fe2c26a53013f3_7_35]]
students.


Changing the value of [[IMAGE:34fe2c26a53013f3_7_36]] will not alter the median score of the [[IMAGE:34fe2c26a53013f3_7_37]] students.


Changing the value of [[IMAGE:34fe2c26a53013f3_7_38]] will definitely change the total range of scores.

Changing the value of [[IMAGE:34fe2c26a53013f3_7_39]] will definitely change the Inter quartile Range (IQR) of
the data set.

Published solution
**1. Given:** \(600\) students. Rahul's score \(x\) is at the \(70^{th}\) percentile. The score is raised slightly but stays below the \(71^{st}\) percentile value.
**2. Concept:** The \(p^{th}\) percentile is the score below which (or at which) about \(p\%\) of the data lies. A small change in one score changes the sum, but changes the median, quartiles or range only if that score crosses their position.
**3. Check each statement:**
- Number at or below \(x\): \(70\%\) of \(600=0.70\times600=420\). True.
- Mean \(=\) (sum of scores)\(/600\). Increasing \(x\) increases the sum, so the mean definitely changes. True.
- Median is the \(50^{th}\) percentile. Rahul is at the \(70^{th}\) percentile and stays below the \(71^{st}\) percentile value, so he stays above the median position. The middle values do not change, so the median is unchanged. True.
- Range \(=\) max \(-\) min. Rahul is at the \(70^{th}\) percentile, so he is neither the maximum nor the minimum. The range is not changed. False.
- IQR \(=Q_3-Q_1\) (\(75^{th}\) and \(25^{th}\) percentiles). Rahul moves from the \(70^{th}\) percentile to still below the \(71^{st}\), so he stays below \(Q_3\). \(Q_1\) and \(Q_3\) are unchanged, so IQR is unchanged. False.
**4. Conclude:** The first three statements are correct.
Answer: A, B, C — about \(420\) students at or below \(x\); the mean definitely changes; the median does not change.
A: **Correct:** \(70\%\) of \(600\) is \(420\) students at or below \(x\).
B: **Correct:** The sum of scores changes when \(x\) changes, so the mean definitely changes.
C: **Correct:** \(x\) stays above the median position, so the middle value is unchanged.
D: **Incorrect:** \(x\) is at the \(70^{th}\) percentile, not the maximum or minimum, so the range does not change.
E: **Incorrect:** \(x\) stays below the \(75^{th}\) percentile, so \(Q_1\) and \(Q_3\) stay the same and IQR does not change.
Question 13 NAT · 4.0 marks
The mean of the [[IMAGE:34fe2c26a53013f3_8_40]] values in a given dataset is [[IMAGE:34fe2c26a53013f3_8_41]] . When these values are rearranged in a
certain order, the mean of the first [[IMAGE:34fe2c26a53013f3_8_42]] values and the last [[IMAGE:34fe2c26a53013f3_8_43]] values are [[IMAGE:34fe2c26a53013f3_8_44]] and [[IMAGE:34fe2c26a53013f3_8_45]] respectively.
Find the tenth value after the rearrangement.






Published solution
**1. Given:** \(30\) values with mean \(100\). The first \(10\) values have mean \(60\). The last \(21\) values have mean \(120\).
**2. Find the sums:**
\[
\text{Total}=30\times100=3000,\quad \text{First }10=10\times60=600,\quad \text{Last }21=21\times120=2520
\]
**3. Notice the overlap:** First \(10\) values and last \(21\) values together count \(10+21=31\) values, but there are only \(30\). So exactly one value is counted twice. It is the value at position \(10\) (the last of the first group, and the first of the last group, since positions \(10\) to \(30\) are the last \(21\)).
**4. Set up the equation:**
\[
600+2520 = 3000 + x_{10}
\]
\[
3120 = 3000 + x_{10}
\]
**5. Solve:**
\[
x_{10}=3120-3000=120
\]
Answer: \(120\).
Question 14 NAT · 2.0 marks
Figure 2 shows a scatter plot of calorie intake ( [[IMAGE:34fe2c26a53013f3_8_46]] ) and weight gain in a week ( [[IMAGE:34fe2c26a53013f3_8_47]] ) for a person.
[[IMAGE:34fe2c26a53013f3_8_48]]
Based on the relationship shown in Figure 2, what is the value of [[IMAGE:34fe2c26a53013f3_8_49]] when [[IMAGE:34fe2c26a53013f3_8_50]] ?





Published solution
**1. Read the points from Figure 2:** \((10,3),\ (20,5),\ (30,7),\ (40,9)\).
**2. Find the relationship:** When \(x\) increases by \(10\), \(y\) increases by \(2\) each time. So the points lie on a straight line.
\[
\text{slope}=\frac{5-3}{20-10}=\frac{2}{10}=0.2
\]
**3. Write the line equation** using the point \((10,3)\):
\[
y-3=0.2(x-10)\ \Rightarrow\ y=0.2x+1
\]
Check with \((40,9)\): \(0.2\times40+1=9\). Correct.
**4. Substitute \(x=100\):**
\[
y=0.2\times100+1=21
\]
Answer: \(21\).
Question 15 NAT · 2.0 marks
A survey was conducted among a group of students to determine their preferred subjects. The
raw frequencies of their responses are displayed in the bar chart in Figure 3.
[[IMAGE:34fe2c26a53013f3_9_51]]
Based on the bar chart, what is the **relative frequency** of students who prefer Subject B? Enter
the answer correct to 2 decimal places.

Published solution
**1. Read the bar chart (Figure 3):** Sub A \(=45\), Sub B \(=75\), Sub C \(=30\), Sub D \(=50\).
**2. Find the total:**
\[
45+75+30+50=200
\]
**3. Formula:** Relative frequency \(=\dfrac{\text{frequency of Subject B}}{\text{total}}\).
**4. Substitute:**
\[
\frac{75}{200}=0.375\approx 0.38
\]
Answer: \(0.38\).
Question 16 NAT · 5.0 marks
Table 2 shows the number of hours spent studying ( [[IMAGE:34fe2c26a53013f3_10_52]] ) and the corresponding test scores ( [[IMAGE:34fe2c26a53013f3_10_53]] ) for
5 students.
[[IMAGE:34fe2c26a53013f3_10_54]]
Calculate the Pearson correlation coefficient between [[IMAGE:34fe2c26a53013f3_10_55]] and [[IMAGE:34fe2c26a53013f3_10_56]] . Enter the answer correct to 2
decimal places.





Published solution
**1. Given (Table 2):** \(X=1,2,3,4,5\) and \(Y=3,4,6,5,7\).
**2. Formula:**
\[
r=\frac{\sum (X-\bar X)(Y-\bar Y)}{\sqrt{\sum (X-\bar X)^2\ \sum (Y-\bar Y)^2}}
\]
**3. Find the means:**
\[
\bar X=\frac{15}{5}=3,\qquad \bar Y=\frac{25}{5}=5
\]
**4. Make the table of deviations:**
\[
\begin{array}{c|c|c|c|c}
X & Y & X-\bar X & Y-\bar Y & (X-\bar X)(Y-\bar Y) \\
\hline
1 & 3 & -2 & -2 & 4 \\
2 & 4 & -1 & -1 & 1 \\
3 & 6 & 0 & 1 & 0 \\
4 & 5 & 1 & 0 & 0 \\
5 & 7 & 2 & 2 & 4
\end{array}
\]
**5. Add up the columns:**
\[
\sum (X-\bar X)(Y-\bar Y)=4+1+0+0+4=9
\]
\[
\sum (X-\bar X)^2=4+1+0+1+4=10
\]
\[
\sum (Y-\bar Y)^2=4+1+1+0+4=10
\]
**6. Substitute:**
\[
r=\frac{9}{\sqrt{10\times10}}=\frac{9}{10}=0.90
\]
Answer: \(0.90\).