Year 11 · Statistical analysis
Quick tips — memory joggers
Data & sampling
Measures of centre
Measures of spread
Outliers, shape & box-plots
Level 1 · Fluency
To find the average height of all Year 11 students in a large school, is it more practical to use a census or a sample?
A sample — measuring every student (a census) is time-consuming, so a representative sample is usually more practical.
To find out how many hours all students in a country sleep, is a census or a sample more practical?
A sample — surveying every student in the country is impractical, so a representative sample is used.
A teacher wants the average mark of her own class of 25 students. Is a census or a sample more practical?
A census — the class is small enough to include every student, so a census is practical and exact.
Level 2 · Application
A company wants to know customer satisfaction across 10 000 customers. Explain why a sample is used and one requirement for it to be reliable.
Surveying all 10 000 is costly and slow, so a sample is used; to be reliable it must be representative (e.g. randomly chosen and large enough).
A manufacturer wants to know the average lifetime of its light globes. Explain why a sample is used and why a census is impossible.
Testing a globe’s lifetime destroys it, so testing every globe (a census) would leave none to sell; a sample gives a practical estimate.
A magazine wants readers’ opinions on a redesign across 50 000 subscribers. Explain why a sample is used and one requirement for it to be reliable.
Surveying all 50 000 is slow and costly, so a sample is used; to be reliable it must be representative (randomly chosen and large enough).
Level 3 · Further Application
A council wants residents’ views on a new park.
(a) A sample — a full census of all residents is impractical, and a well-chosen sample can represent the population.
(b) Select residents randomly (or use stratified sampling across age groups/suburbs) so all groups are fairly represented.
A supermarket chain wants customers’ views on opening hours across 40 stores.
(a) A sample — surveying every customer at all 40 stores is impractical, and a well-chosen sample can represent them.
(b) Randomly select customers across different stores and times of day (stratified) so all types of shopper are included.
A university wants feedback from its 20 000 students on a new library.
(a) A sample — contacting all 20 000 students is costly and slow, and a representative sample can stand in for them.
(b) Use stratified sampling across faculties and year groups so each group is fairly represented.
Level 1 · Fluency
Selecting every 10th person on a list is which sampling method?
Systematic sampling.
Names are drawn from a hat to select participants. Which sampling method is this?
Random (simple random) sampling.
A population is split into age groups and members are chosen from each group in proportion. Which sampling method is this?
Stratified sampling.
Level 2 · Application
A survey is advertised online and people choose whether to respond. Name this sampling method and one problem with it.
Self-selected sampling. Problem: respondents may not represent the population (those with strong opinions are over-represced), causing bias.
To estimate the number of fish in a lake, scientists tag 50 fish, release them, and later check what fraction of a new catch is tagged. Name this method and one assumption it relies on.
Capture–recapture. It assumes the tagged fish mix evenly back into the population and that the population does not change between catches.
A factory tests every 20th item coming off a production line. Name this method and one problem it can have.
Systematic sampling. Problem: if faults occur in a regular cycle matching the step, the sample can be biased.
Level 3 · Further Application
A school of 800 students has 500 juniors and 300 seniors. A sample of 80 is taken with 50 juniors and 30 seniors.
(a) Stratified sampling.
(b) School: \(\dfrac{500}{800} = 62.5\%\) junior; sample: \(\dfrac{50}{80} = 62.5\%\) junior — the same proportion, so it is stratified correctly.
A company has 1200 employees: 900 full-time and 300 part-time. A sample of 60 is taken with 45 full-time and 15 part-time.
(a) Stratified sampling.
(b) Company: \(\dfrac{900}{1200} = 75\%\) full-time; sample: \(\dfrac{45}{60} = 75\%\) full-time — the same proportion, so it is stratified correctly.
A club has 250 juniors and 150 seniors. A stratified sample of 40 is required.
25 juniors
(a) Stratified sampling.
(b) Juniors are \(\dfrac{250}{400} = 62.5\%\) of the club, so \(0.625 \times 40 =\) 25 juniors.
Level 1 · Fluency
State one advantage of random sampling.
Every member has an equal chance of selection, which reduces bias.
State one disadvantage of a census compared with a sample.
A census is time-consuming and expensive because every member of the population must be measured.
State one advantage of systematic sampling.
It is quick and easy to carry out — you just select every \(k\)th member from a list.
Level 2 · Application
Compare systematic sampling with simple random sampling in terms of ease and bias.
Systematic is easier to carry out (just count in a fixed step) but can be biased if the list has a repeating pattern; random sampling avoids that pattern bias but needs a full list and a random selector.
Compare stratified sampling with simple random sampling in terms of representing subgroups.
Stratified sampling guarantees each subgroup appears in proportion, so subgroups are always represented; simple random sampling may by chance under-represent a small subgroup.
Compare a self-selected sample with a random sample in terms of bias.
A self-selected sample is prone to bias because only people with strong views or interest respond; a random sample gives everyone an equal chance and so reduces bias.
Level 3 · Further Application
A researcher must survey a large, diverse population.
(a) Advantage: guarantees each subgroup is represented in proportion. Disadvantage: you must know the subgroup sizes and it is more work to organise.
(b) When the population has clear subgroups that differ (e.g. age or region) and you want each fairly represented.
A researcher is choosing between systematic and random sampling for a long membership list.
(a) Advantage: fast and simple to apply from an ordered list. Disadvantage: it needs a full list and can be biased if the list has a repeating pattern.
(b) When the list repeats in a cycle that matches the sampling step (e.g. every \(k\)th house is a corner block).
A survey will be run in a large workplace.
(a) Advantage: it is exact and includes everyone, with no sampling error. Disadvantage: it is costly and time-consuming.
(b) When the population is small, or when complete accuracy is essential.
Level 1 · Fluency
A survey on exercise is conducted only at a gym. Why might the results be biased?
People at a gym exercise more than the general population, so the sample is not representative.
A survey about reading habits is conducted only in a library. Why might the results be biased?
People in a library read more than the general population, so the sample is not representative.
A survey about phone use is conducted only among teenagers at a shopping centre. Why might the results be biased?
Teenagers use phones differently from other age groups, so the sample does not represent the whole population.
Level 2 · Application
A TV station asks viewers to phone in to vote on an issue. Identify two reasons the results may be unreliable.
It is self-selected (only motivated viewers respond) and limited to that station’s audience, so it does not represent the wider population; people can also vote multiple times.
An online poll on a gaming website asks whether video games should be cheaper. Identify two reasons the results may be unreliable.
It is self-selected (only interested visitors respond) and limited to gamers, so it is not representative; people may also vote more than once.
A researcher surveys shoppers at 9am on a weekday about employment. Identify two reasons the results may be biased.
At that time many respondents are not in full-time work (e.g. retirees, shift workers), so the sample is unrepresentative, and it is taken at one location only.
Level 3 · Further Application
A survey asks: “Do you agree that our excellent library should get more funding?”
(a) It is a leading question — the word “excellent” pushes respondents toward agreeing, biasing the result.
(b) e.g. “Should the library receive more funding?”
A survey asks: “Do you support the sensible plan to reduce wasteful council spending?”
(a) It is a leading question — the words “sensible” and “wasteful” push respondents toward agreeing.
(b) e.g. “Do you support the council’s plan to reduce spending?”
A survey asks: “How much do you love our award-winning coffee?”
(a) It is leading and assumes the respondent already loves the coffee (“award-winning”, “love”), biasing answers positively.
(b) e.g. “How would you rate our coffee?”
Level 1 · Fluency
Is “eye colour” categorical or numerical?
Categorical.
Is “number of pets” categorical or numerical?
Numerical.
Is “country of birth” categorical or numerical?
Categorical.
Level 2 · Application
Classify each as categorical or numerical: (i) height in cm, (ii) favourite sport, (iii) number of siblings.
(i) numerical, (ii) categorical, (iii) numerical.
Classify each as categorical or numerical: (i) shoe size, (ii) hair colour, (iii) mass in kg.
(i) numerical, (ii) categorical, (iii) numerical.
Classify each as categorical or numerical: (i) car colour, (ii) daily temperature, (iii) number of text messages sent.
(i) categorical, (ii) numerical, (iii) numerical.
Level 3 · Further Application
A survey records postcode, temperature and satisfaction rating (1–5 stars).
(a) Postcode: categorical; temperature: numerical; star rating: categorical (ordinal).
(b) A postcode is a label for an area — arithmetic on it (e.g. averaging) is meaningless, so it is categorical despite looking numeric.
A survey records phone number, age and gender.
(a) Phone number: categorical; age: numerical; gender: categorical.
(b) A phone number is a label/identifier — arithmetic on it (e.g. averaging) is meaningless — so it is categorical despite being written as digits.
A dataset records jersey number, height and team name.
(a) Jersey number: categorical; height: numerical; team name: categorical.
(b) A jersey number just labels a player; averaging or ordering it has no meaning, so it is categorical.
Level 1 · Fluency
Is “t-shirt size (S, M, L)” ordinal or nominal?
Ordinal — the categories have a natural order.
Is “country of birth” ordinal or nominal?
Nominal — the categories have no natural order.
Is “medal won (gold, silver, bronze)” ordinal or nominal?
Ordinal — the categories have a natural order.
Level 2 · Application
Classify each as ordinal or nominal: (i) blood type, (ii) exam grade (A–E).
(i) nominal (no order); (ii) ordinal (A–E has an order).
Classify each as ordinal or nominal: (i) eye colour, (ii) survey response (disagree, neutral, agree).
(i) nominal (no order); (ii) ordinal (the responses have an order).
Classify each as ordinal or nominal: (i) star rating (1–5), (ii) type of pet.
(i) ordinal (ratings have an order); (ii) nominal (no order).
Level 3 · Further Application
A survey records hair colour and a movie rating of “poor / good / excellent”.
(a) Hair colour: nominal; movie rating: ordinal.
(b) Ordinal categories have a meaningful order (rankings), while nominal categories are just names/labels with no order.
A survey records nationality and a fitness level of “low / medium / high”.
(a) Nationality: nominal; fitness level: ordinal.
(b) Ordinal categories have a meaningful order (rankings), while nominal categories are just names/labels with no order.
A survey records favourite colour and a satisfaction level of “unhappy / okay / happy”.
(a) Favourite colour: nominal; satisfaction level: ordinal.
(b) A postcode only labels an area and has no meaningful order or arithmetic, so it is nominal.
Level 1 · Fluency
Is “number of cars in a car park” discrete or continuous?
Discrete (it can only take whole-number values).
Is “the height of a tree” discrete or continuous?
Continuous (it can take any value in a range).
Is “the number of students in a class” discrete or continuous?
Discrete (whole-number values only).
Level 2 · Application
Classify each as discrete or continuous: (i) time to run 100 m, (ii) number of goals scored.
(i) continuous, (ii) discrete.
Classify each as discrete or continuous: (i) mass of a baby, (ii) number of eggs in a nest.
(i) continuous, (ii) discrete.
Classify each as discrete or continuous: (i) number of cars sold, (ii) length of a phone call.
(i) discrete, (ii) continuous.
Level 3 · Further Application
A dataset records the mass of parcels and the number of items in each parcel.
(a) Mass: continuous; number of items: discrete.
(b) Continuous data can take any value in a range (measured); discrete data can take only separate, usually whole-number, values (counted).
A dataset records the volume of drink in each bottle and the number of bottles in each crate.
(a) Volume: continuous; number of bottles: discrete.
(b) Continuous data is measured and can take any value in a range; discrete data is counted and takes only separate whole-number values.
A dataset records the time each runner takes and the number of runners in each heat.
(a) Time: continuous; number of runners: discrete.
(b) e.g. continuous: a person’s height; discrete: the number of siblings a person has.
Level 1 · Fluency
Which display is suitable for showing the proportion of each category in a whole — a bar chart or a stem-and-leaf plot?
A bar chart (or a sector/pie graph) — stem-and-leaf is for numerical data.
Which display best shows how a single total is divided among a few categories — a pie chart or a stem-and-leaf plot?
A pie (sector) chart — it shows each category as part of the whole.
Which display is suitable for showing a numerical distribution grouped into class intervals — a histogram or a pie chart?
A histogram.
Level 2 · Application
For a dataset of the number of pets owned by 30 households, name a suitable graph and justify.
A dot plot or column graph — the data is discrete numerical, and these show the frequency of each value clearly.
For continuous data on the daily rainfall (mm) over 60 days, name a suitable graph and justify.
A histogram — the data is continuous numerical, so it can be grouped into class intervals and shown with adjoining bars.
For categorical data on students’ favourite subjects, name a suitable graph and justify.
A bar chart (or Pareto chart) — categorical data has no numerical scale, so separated bars show each category’s frequency clearly.
Level 3 · Further Application
A researcher has (i) categorical data on preferred transport and (ii) continuous data on commute times.
(a) (i) bar chart / Pareto chart; (ii) histogram.
(b) A histogram groups numerical data into continuous class intervals; transport preference has no numerical scale to group, so a bar chart (with gaps) is used instead.
A researcher has (i) categorical data on favourite sport and (ii) discrete data on the number of goals scored per match.
(a) (i) bar chart / Pareto chart; (ii) dot plot or column graph.
(b) A dot plot places dots on a number line, which needs a numerical scale — favourite sport has no numerical scale, so a bar chart is used instead.
A researcher has (i) categorical data on eye colour and (ii) continuous data on students’ heights.
(a) (i) bar chart / pie chart; (ii) histogram.
(b) A histogram groups numerical values into continuous intervals; eye colour has no numerical scale to group, so a bar chart (with gaps) is used.
Level 1 · Fluency
In a Pareto chart, the bars are arranged in what order?
From tallest to shortest (most frequent category first).
In a Pareto chart, is the cumulative percentage line increasing or decreasing from left to right?
Increasing — it is a running total, so it rises to 100%.
What type of data is displayed with a Pareto chart — categorical or continuous numerical?
Categorical data.
Level 2 · Application
A Pareto chart shows complaint types with a cumulative line. Explain what the cumulative line shows.
It shows the running total (cumulative frequency or percentage), so you can see how much of the total is accounted for by the leading categories.
A Pareto chart of machine faults helps a manager decide where to focus repairs. Explain how the chart helps.
The tallest bars (most frequent faults) are on the left, and the cumulative line shows they account for most of the problems, so the manager can target the few faults causing the majority of issues.
Explain the difference between a bar chart and a Pareto chart.
A Pareto chart is a bar chart whose bars are ordered from most to least frequent and includes a cumulative percentage line; an ordinary bar chart can be in any order and has no cumulative line.
Level 3 · Further Application
Complaints: Late 40, Wrong item 25, Damaged 20, Other 15.
65%
(a) Late, Wrong item, Damaged, Other.
(b) Total \(= 100\); first two \(= 40 + 25 = 65\), so 65%.
Machine faults: Jam 30, Overheat 18, Leak 8, Other 4.
80%
(a) Jam, Overheat, Leak, Other.
(b) Total \(= 60\); first two \(= 30 + 18 = 48\), so \(\dfrac{48}{60} =\) 80%.
Favourite fruit: Apple 24, Banana 15, Orange 6, Other 5.
78%
(a) Apple (the most frequent category).
(b) Total \(= 50\); first two \(= 24 + 15 = 39\), so \(\dfrac{39}{50} =\) 78%.
Level 1 · Fluency
In a stem-and-leaf plot, the value 24 is recorded with stem 2 and leaf what?
Leaf 4.
In a stem-and-leaf plot, the value 57 is recorded with leaf 7 and which stem?
Stem 5.
In a stem-and-leaf plot with stem 3 and leaf 6, what value is represented?
36.
Level 2 · Application
Given the data 12, 15, 15, 18, 21, 23, 23, 29, construct a stem-and-leaf plot.
Stem 1 | 2 5 5 8 Stem 2 | 1 3 3 9.
Given the data 31, 34, 34, 38, 42, 45, 45, 49, construct a stem-and-leaf plot.
Stem 3 | 1 4 4 8 Stem 4 | 2 5 5 9.
Given the data 6, 9, 11, 14, 14, 20, 22, construct a stem-and-leaf plot (stems as tens).
Stem 0 | 6 9 Stem 1 | 1 4 4 Stem 2 | 0 2.
Level 3 · Further Application
A stem-and-leaf plot shows the runs scored in ten innings: stem 1 | 4 5, stem 2 | 1 3 3 3 8, stem 3 | 0 2 4.
23; 4 innings
(a) The repeated value is 23 (stem 2, leaf 3 appears three times), so the mode is 23.
(b) Values above 25: 28, 30, 32, 34 → 4 innings.
A stem-and-leaf plot shows test marks: stem 4 | 2 5, stem 5 | 1 3 3 3 7, stem 6 | 0 4.
53; 3 students
(a) The value 53 appears three times (stem 5, leaf 3), so the mode is 53.
(b) Marks above 55: 57, 60, 64 → 3 students.
A stem-and-leaf plot shows house ages (years): stem 1 | 2 8, stem 2 | 0 4 4 4 9, stem 3 | 1 5.
24; 6 houses
(a) The value 24 appears three times, so the mode is 24.
(b) Ages below 25: 12, 18, 20, 24, 24, 24 → 6 houses.
Level 1 · Fluency
A frequency table has a value with frequency 6. What does that frequency mean?
That value occurred 6 times in the dataset.
In a frequency table, a value has frequency 0. What does that mean?
That value did not occur in the dataset.
In a relative frequency table, a value has relative frequency 0.25. What does that mean?
That value made up 25% (a quarter) of the data.
Level 2 · Application
A cumulative frequency column reaches 20 at the value “3 pets”. Interpret this.
20 households own 3 or fewer pets.
A cumulative frequency column reaches 35 at the value “2 goals”. Interpret this.
35 matches had 2 or fewer goals.
In a frequency table, “number of siblings = 1” has frequency 12 out of 40 students. Find the relative frequency.
0.3 (30%)
\(\dfrac{12}{40} = 0.3\), i.e. 30%.
Level 3 · Further Application
A frequency table of daily downloads has frequencies 3, 5, 8, 4 for 0, 1, 2, 3 downloads.
1.65
(a) \(3 + 5 + 8 + 4 = 20\) days.
(b) Total downloads \(= 0\times3 + 1\times5 + 2\times8 + 3\times4 = 33\); mean \(= 33 \div 20 =\) 1.65.
A frequency table of goals per match has frequencies 4, 6, 7, 3 for 0, 1, 2, 3 goals.
1.45
(a) \(4 + 6 + 7 + 3 = 20\) matches.
(b) Total goals \(= 0\times4 + 1\times6 + 2\times7 + 3\times3 = 29\); mean \(= 29 \div 20 =\) 1.45.
A frequency table of pets per household has frequencies 5, 9, 4, 2 for 0, 1, 2, 3 pets.
1.15
(a) \(5 + 9 + 4 + 2 = 20\) households.
(b) Total pets \(= 0\times5 + 1\times9 + 2\times4 + 3\times2 = 23\); mean \(= 23 \div 20 =\) 1.15.
Level 1 · Fluency
To fairly compare two datasets of different sizes in a table, should you compare frequencies or relative frequencies?
Relative frequencies (proportions), so different totals can be compared fairly.
When comparing two datasets on the same graph, why is it important to use the same scale on both axes?
So the comparison is fair — a different scale could make one dataset look larger or more spread out than it really is.
To compare the shape of two datasets of very different sizes, should you compare counts or percentages?
Percentages (relative frequencies), so the different totals do not distort the comparison.
Level 2 · Application
Two classes sit the same test. Class A’s median is 68 and Class B’s median is 74. Interpret this comparison.
Class B’s typical (middle) score is higher, so on average Class B performed better on the test.
Two teams’ scores have means: Team A 45, Team B 52. Interpret this comparison.
Team B has a higher average score, so on average Team B performed better.
Two workers’ assembly times have medians: Worker X 12 min, Worker Y 9 min. Interpret this comparison.
Worker Y’s typical time is shorter, so Worker Y is generally faster at assembling.
Level 3 · Further Application
Two shops record daily sales for a week. Shop X: median $500, range $300. Shop Y: median $520, range $700.
(a) Shop Y (higher median).
(b) Shop X is more consistent — its range ($300) is smaller than Shop Y’s ($700), so its daily sales vary less.
Two machines’ fill weights: Machine P median 500 g, range 20 g; Machine Q median 502 g, range 60 g.
(a) Machine Q (higher median).
(b) Machine P is more consistent — its range (20 g) is smaller than Machine Q’s (60 g), so its fills vary less.
Two bus routes’ journey times: Route 1 median 30 min, IQR 4 min; Route 2 median 28 min, IQR 12 min.
(a) Route 2 (lower median time).
(b) Route 1 is more reliable — its IQR (4 min) is much smaller than Route 2’s (12 min), so its times are more consistent.
Level 1 · Fluency
Which display is designed specifically to compare the spread and centre of two datasets side by side?
Parallel (side-by-side) box-plots.
Which display uses two sets of leaves on either side of a shared stem to compare two datasets?
A back-to-back stem-and-leaf plot.
To compare the frequencies of the same categories for two groups, which display is suitable?
A side-by-side (grouped) column/bar chart.
Level 2 · Application
Explain why parallel box-plots are useful for comparing two datasets.
They show the median, quartiles, range and any outliers of both datasets on the same scale, making differences in centre and spread easy to see.
Explain why a back-to-back stem-and-leaf plot is useful for comparing two small datasets.
It lines up both datasets against the same stems, so you can compare their shape, centre and spread while still seeing every individual value.
Explain why relative frequencies (not raw counts) are used when comparing two datasets of different sizes in a table.
Relative frequencies convert counts to proportions of each total, so datasets with different totals can be compared fairly.
Level 3 · Further Application
A researcher wants to compare exam marks of two large classes.
(a) Parallel box-plots — they align both distributions on one scale for direct comparison of centre and spread.
(b) The median (or IQR / range).
A researcher wants to compare the daily temperatures of two cities over a year.
(a) Parallel box-plots — they place both distributions on one scale so centre and spread can be compared directly.
(b) The median (or IQR / range).
A shop wants to compare the sizes of items sold in two branches.
(a) A side-by-side column graph (or parallel box-plots) — it shows both branches on the same axes.
(b) Placing both on one scale makes differences easy to see at a glance, avoiding mismatched scales.
Level 1 · Fluency
For showing how a quantity changes over 12 months, is a line graph or a pie chart more suitable?
A line graph — it shows change over time.
To show the proportion of a class that chose each of four options, is a pie chart or a line graph more suitable?
A pie chart — it shows each option as part of the whole.
To compare the population of five different towns, is a bar chart or a line graph more suitable?
A bar chart — it compares separate categories.
Level 2 · Application
A council wants to show the proportion of its budget spent on each service. Recommend a display and justify.
A sector (pie) graph or a 100% stacked bar — it shows each part as a proportion of the whole budget.
A scientist wants to show how a plant’s height changed each week over 10 weeks. Recommend a display and justify.
A line graph — it shows a continuous change over time, revealing the growth trend.
A school wants to compare the number of students in each of six houses. Recommend a display and justify.
A bar/column chart — the houses are separate categories, so bars let you compare their sizes directly.
Level 3 · Further Application
A business has monthly revenue for two years and wants to highlight the trend and any seasonal pattern.
(a) A line graph over time — it reveals the overall trend and repeating seasonal peaks.
(b) A pie chart shows proportions at a single point, not change over time, so it cannot show a trend.
A weather station has daily maximum temperatures for two years and wants to show the yearly trend and seasonal pattern.
(a) A line graph over time — it reveals the overall trend and repeating seasonal peaks.
(b) A pie chart shows proportions at one moment, not change over time, so it cannot show a trend.
A council has the percentage of its budget spent on five services this year.
(a) A pie chart (or 100% stacked bar) — it shows each service as a proportion of the whole budget.
(b) A line graph is for change over time; these are categories at one point in time, so joining them with a line has no meaning.
Level 1 · Fluency
What is a population in statistics?
The entire group being studied.
What is a sample in statistics?
A subset (part) of the population that is actually studied.
Is measuring every member of a group a census or a sample?
A census.
Level 2 · Application
A researcher measures 200 of a city’s 50 000 households. Identify the population and the sample.
Population: all 50 000 households; sample: the 200 measured.
A pollster surveys 1500 of a country’s 16 million voters. Identify the population and the sample.
Population: all 16 million voters; sample: the 1500 surveyed.
A biologist tags 80 of the estimated 5000 fish in a lake. Identify the population and the sample.
Population: all 5000 fish; sample: the 80 tagged fish.
Level 3 · Further Application
A factory checks 60 of the 12 000 items made in a day.
(a) Population: 12 000 items; sample: 60 items.
(b) Testing all items is costly and slow (and may be destructive), so a sample gives a practical estimate of quality.
A bakery checks 40 of the 5000 loaves baked in a day.
(a) Population: 5000 loaves; sample: 40 loaves.
(b) Checking every loaf is slow and may be wasteful (cutting them open), so a sample gives a practical estimate of quality.
A phone maker tests 25 of the 8000 phones assembled in a shift.
(a) Population: 8000 phones; sample: 25 phones.
(b) Testing all 8000 would be costly and time-consuming (and some tests may damage the phone), so a sample estimates quality efficiently.
Level 1 · Fluency
Which symbol represents the population mean — \(\mu\) or \(\bar{x}\)?
\(\mu\) (mu) is the population mean; \(\bar{x}\) is the sample mean.
Which symbol represents the sample mean — \(\mu\) or \(\bar{x}\)?
\(\bar{x}\) is the sample mean; \(\mu\) is the population mean.
Which symbol represents the population standard deviation — \(\sigma\) or \(s\)?
\(\sigma\) (sigma) is the population standard deviation; \(s\) is the sample standard deviation.
Level 2 · Application
State the symbol for the population standard deviation and the sample mean.
Population standard deviation: \(\sigma\) (sigma); sample mean: \(\bar{x}\).
State the symbol for the sample standard deviation and the population mean.
Sample standard deviation: \(s\); population mean: \(\mu\) (mu).
A report uses \(\bar{x}\) and \(s\). Do these describe a population or a sample?
A sample — \(\bar{x}\) and \(s\) are the sample mean and sample standard deviation.
Level 3 · Further Application
A study reports \(\bar{x} = 172\) cm from a sample and estimates \(\mu\) for the population.
(a) \(\bar{x} = 172\) cm is the mean of the sample actually measured; \(\mu\) is the (usually unknown) mean of the whole population.
(b) Measuring the whole population is impractical, so the sample mean \(\bar{x}\) is used as the best estimate of \(\mu\).
A survey finds a sample mean mass \(\bar{x} = 68\) kg and uses it to estimate the population mean \(\mu\).
(a) \(\bar{x} = 68\) kg is the mean of the sample measured; \(\mu\) is the (usually unknown) mean of the whole population.
(b) Measuring the whole population is impractical, so the sample mean \(\bar{x}\) is the best available estimate of \(\mu\).
A quality report gives sample standard deviation \(s = 3\) and estimates the population value \(\sigma\).
(a) \(s = 3\) is the standard deviation of the sample; \(\sigma\) is the standard deviation of the whole population.
(b) Different symbols make it clear whether a statistic comes from a sample (an estimate) or the entire population (the true value).
Level 1 · Fluency
Name one measure of centre and one measure of spread.
Centre: mean (or median). Spread: range (or IQR / standard deviation).
Name two different measures of centre.
Mean and median (the mode is a third).
Name two different measures of spread.
Range and interquartile range (standard deviation is a third).
Level 2 · Application
For the data 4, 7, 8, 9, 12, calculate the mean and the range.
Mean \(= 40 \div 5 = 8\); range \(= 12 - 4 = 8\).
For the data 5, 6, 9, 10, 15, calculate the mean and the range.
Mean \(= 45 \div 5 = 9\); range \(= 15 - 5 = 10\).
For the data 10, 14, 15, 18, 23, calculate the mean and the range.
Mean \(= 80 \div 5 = 16\); range \(= 23 - 10 = 13\).
Level 3 · Further Application
For the data 4, 7, 8, 9, 12:
2
(a) Median \(= 8\); range \(= 8\).
(b) \(Q_1 = 7\) (median of 4, 7), \(Q_3 = 9\) (median of 9, 12); IQR \(= 9 - 7 =\) 2.
For the data 6, 9, 11, 14, 20:
9.5
(a) Median \(= 11\); range \(= 20 - 6 = 14\).
(b) \(Q_1 = 7.5\) (median of 6, 9), \(Q_3 = 17\) (median of 14, 20); IQR \(= 17 - 7.5 =\) 9.5.
For the data 3, 8, 10, 15, 21:
12.5
(a) Median \(= 10\); range \(= 21 - 3 = 18\).
(b) \(Q_1 = 5.5\) (median of 3, 8), \(Q_3 = 18\) (median of 15, 21); IQR \(= 18 - 5.5 =\) 12.5.
Level 1 · Fluency
Find the mode of 3, 5, 5, 7, 9.
5
5 (the most frequent value).
Find the mode of 2, 4, 4, 4, 8, 10.
4
4 occurs three times, more than any other value.
Find the mode of 11, 13, 13, 15, 17, 17, 17.
17
17 appears three times, the most of any value.
Level 2 · Application
A dataset is 2, 2, 4, 6, 6, 9. Explain why it has two modes and name them.
Both 2 and 6 occur twice (more than any other value), so the data is bimodal with modes 2 and 6.
A dataset is 5, 5, 7, 8, 8, 12. Explain why it has two modes and name them.
Both 5 and 8 occur twice (more than any other value), so the data is bimodal with modes 5 and 8.
A dataset is 3, 4, 4, 7, 9, 9. State the modes and the name given to a two-mode dataset.
Modes are 4 and 9; a dataset with two modes is called bimodal.
Level 3 · Further Application
A dataset of shoe sizes sold is 7, 8, 8, 9, 9, 9, 10.
(a) Mode = 9.
(b) The mode is the most commonly sold size, so it tells the shop which size to stock most of — more useful here than a mean “average size”.
A café records drink sizes sold: S, M, M, L, L, L, L, M.
(a) Mode = Large (L).
(b) The mode is the most frequently sold size, so it tells the café which cup size to stock most of — more useful than an “average size”, which is meaningless for categories.
A shop records T-shirt sizes sold: 8, 10, 10, 12, 12, 12, 14.
(a) Mode = 12.
(b) The mode is the most popular size to stock, whereas the mean size (a decimal) would not correspond to a real size the shop sells.
Level 1 · Fluency
Find the mean of 6, 8, 10, 12.
9
\((6 + 8 + 10 + 12) \div 4 =\) 9.
Find the mean of 4, 9, 11, 16.
10
\((4 + 9 + 11 + 16) \div 4 = 40 \div 4 =\) 10.
Find the mean of 5, 7, 12, 20.
11
\((5 + 7 + 12 + 20) \div 4 = 44 \div 4 =\) 11.
Level 2 · Application
For the data 3, 7, 8, 11, 16, calculate the mean and the median.
Mean \(= 45 \div 5 = 9\); median = middle value \(= 8\).
For the data 5, 6, 9, 13, 17, calculate the mean and the median.
Mean \(= 50 \div 5 = 10\); median = middle value \(= 9\).
For the data 2, 4, 7, 8, 9, 12, calculate the mean and the median.
Mean \(= 42 \div 6 = 7\); median \(= \dfrac{7 + 8}{2} = 7.5\).
Level 3 · Further Application
Consider the data 3, 4, 8, 9, 16. A new value \(x\) (with \(x > 16\)) is added.
32
(a) Old sum \(= 40\), so new mean \(= \dfrac{40 + x}{6}\).
(b) Old mean 8, old median 8; new median (of 6 values) \(= \dfrac{8 + 9}{2} = 8.5\). So \(\dfrac{40 + x}{6} - 8 = 8(8.5 - 8) = 4 \Rightarrow \dfrac{40 + x}{6} = 12 \Rightarrow 40 + x = 72 \Rightarrow x =\) 32.
Consider the data 2, 5, 9, 10, 14. A new value \(x\) (with \(x > 14\)) is added.
32
(a) Old sum \(= 40\), so new mean \(= \dfrac{40 + x}{6}\).
(b) \(\dfrac{40 + x}{6} = 12 \Rightarrow 40 + x = 72 \Rightarrow x =\) 32.
The mean of 6, 8, 11, 15 and one more value \(x\) is 12.
20
(a) \(\dfrac{6 + 8 + 11 + 15 + x}{5} = 12\), i.e. \(\dfrac{40 + x}{5} = 12\).
(b) \(40 + x = 60 \Rightarrow x =\) 20.
Level 1 · Fluency
If a dataset has a very large outlier, is the mean or the median more affected?
The mean (it is pulled toward the outlier); the median is resistant.
A single very small outlier is added to a dataset. Which is more affected — the mean or the median?
The mean (it is pulled toward the outlier); the median is resistant.
For a symmetric dataset with no outliers, are the mean and median usually close together or far apart?
Close together (roughly equal).
Level 2 · Application
House prices in a suburb are mostly $600k–$800k but one sold for $5 million. State whether the mean or median better represents a “typical” price, and why.
The median — the single very high sale inflates the mean, so the median better reflects a typical price.
Most workers at a firm earn $50k–$70k, but the CEO earns $2 million. State whether the mean or median better represents a typical wage, and why.
The median — the CEO’s very high salary inflates the mean, so the median better reflects a typical wage.
In a class, most test marks are 60–80 but two students scored 5. State whether the mean or median better represents typical performance, and why.
The median — the two very low marks drag the mean down, so the median better represents the typical mark.
Level 3 · Further Application
Weekly wages at a firm are mostly $900–$1100, but the owner earns $12 000.
(a) The median.
(b) The owner’s very high wage is an outlier that pulls the mean upward, making it unrepresentative; the median is not affected by that extreme value.
House prices on a street are mostly $700k–$900k, but one mansion sold for $6 million.
(a) The median.
(b) The $6 million sale is an outlier that pulls the mean upward, making it unrepresentative; the median is not affected by that single extreme value.
Donation amounts at a fundraiser are mostly $10–$50, but one donor gave $10 000.
(a) The median.
(b) The single $10 000 donation is an outlier that greatly inflates the mean; the median ignores that extreme and better reflects a typical donation.
Level 1 · Fluency
Find the range of 12, 18, 25, 30, 41.
29
\(41 - 12 =\) 29.
Find the range of 8, 15, 22, 27, 34.
26
\(34 - 8 =\) 26.
Find the range of 45, 52, 60, 71, 88.
43
\(88 - 45 =\) 43.
Level 2 · Application
For the data 3, 5, 7, 8, 12, 15, 18, 20, find \(Q_1\), \(Q_3\) and the IQR.
10.5
Lower half 3,5,7,8 → \(Q_1 = 6\); upper half 12,15,18,20 → \(Q_3 = 16.5\); IQR \(= 16.5 - 6 =\) 10.5.
For the data 4, 6, 9, 10, 13, 17, 19, 24, find \(Q_1\), \(Q_3\) and the IQR.
10.5
Lower half 4, 6, 9, 10 → \(Q_1 = 7.5\); upper half 13, 17, 19, 24 → \(Q_3 = 18\); IQR \(= 18 - 7.5 =\) 10.5.
For the data 2, 4, 5, 8, 10, 11, find \(Q_1\), \(Q_3\) and the IQR.
6
Lower half 2, 4, 5 → \(Q_1 = 4\); upper half 8, 10, 11 → \(Q_3 = 10\); IQR \(= 10 - 4 =\) 6.
Level 3 · Further Application
A dataset has mean 50 and standard deviation 8 (from technology).
(a) It measures the typical spread of values about the mean — on average, values differ from 50 by about 8.
(b) \((66 - 50) \div 8 = 2\) standard deviations above the mean.
A dataset has mean 60 and standard deviation 5 (from technology).
(a) It measures the typical spread of values about the mean — on average values differ from 60 by about 5.
(b) \((72 - 60) \div 5 = 2.4\) standard deviations above the mean.
A dataset has mean 100 and standard deviation 12.
(a) On average, values differ from the mean of 100 by about 12 — it measures the spread.
(b) \((100 - 76) \div 12 = 2\) standard deviations below the mean.
Level 1 · Fluency
Adding a very large outlier to a dataset will increase which — the mean or the mode?
The mean (the mode is usually unchanged).
Removing a very large outlier from a dataset will decrease which — the mean or the mode?
The mean (the mode is usually unchanged).
Adding a very large outlier to a dataset has the biggest effect on the range or the median?
The range (it depends on the extremes); the median barely changes.
Level 2 · Application
The data 4, 6, 7, 8 has mean 6.25. A value of 40 is added. State the effect on the mean and the range.
New mean \(= (25 + 40) \div 5 = 13\) (much higher); range jumps from 4 to 36 — both increase greatly.
The data 5, 6, 8, 9 has mean 7. A value of 32 is added. State the effect on the mean and the range.
New mean \(= (28 + 32) \div 5 = 12\) (much higher); range rises from 4 to 27 — both increase greatly.
The data 10, 12, 13, 15 has mean 12.5. A value of 60 is added. State the effect on the mean and the median.
New mean \(= (50 + 60) \div 5 = 22\) (much higher, pulled up by 60); the median only rises from 12.5 to 13 — barely affected.
Level 3 · Further Application
A dataset of 9 values has mean 20 and median 19.
(a) The mean rises sharply (the total gains 200 over 10 values), while the median barely changes (it stays near the middle of the ordered data).
(b) The median — it resists the outlier and better represents the typical value.
A dataset of 7 values has mean 15 and median 14.
(a) The mean rises sharply (the total gains 120 over 8 values), while the median barely changes (it stays near the middle of the ordered data).
(b) The median — it resists the outlier and better represents the typical value.
A dataset of 11 values has mean 30 and median 28.
(a) The mean falls noticeably (the low value reduces the total), while the median moves only slightly.
(b) The median — it is not distorted by the single low outlier.
Level 1 · Fluency
The upper outlier boundary is \(Q_3 + 1.5 \times IQR\). If \(Q_3 = 10\) and IQR \(= 6\), find this boundary.
19
\(10 + 1.5 \times 6 =\) 19.
The lower outlier boundary is \(Q_1 - 1.5 \times IQR\). If \(Q_1 = 20\) and IQR \(= 8\), find this boundary.
8
\(20 - 1.5 \times 8 = 20 - 12 =\) 8.
If \(Q_3 = 45\) and IQR \(= 10\), find the upper outlier boundary \(Q_3 + 1.5 \times IQR\).
60
\(45 + 1.5 \times 10 = 45 + 15 =\) 60.
Level 2 · Application
A five-number summary is 2, 5, 8, 14, 30. Determine whether 30 is an outlier.
30 is an outlier
IQR \(= 14 - 5 = 9\); upper boundary \(= 14 + 1.5 \times 9 = 27.5\). Since \(30 > 27.5\), 30 is an outlier.
A five-number summary is 5, 10, 14, 18, 40. Determine whether 40 is an outlier.
40 is an outlier
IQR \(= 18 - 10 = 8\); upper boundary \(= 18 + 1.5 \times 8 = 30\). Since \(40 > 30\), 40 is an outlier.
A five-number summary is 1, 12, 18, 20, 27. Determine whether the minimum, 1, is an outlier.
1 is not an outlier
IQR \(= 20 - 12 = 8\); lower boundary \(= 12 - 1.5 \times 8 = 0\). Since \(1 > 0\), it is not below the boundary, so 1 is not an outlier.
Level 3 · Further Application
A dataset has \(Q_1 = 22\), \(Q_3 = 34\).
not
(a) IQR \(= 12\); lower \(= 22 - 1.5 \times 12 = 4\); upper \(= 34 + 1.5 \times 12 = 52\).
(b) 4 is exactly on the lower boundary, so it is not beyond it — it is not an outlier (values must be below 4 to qualify).
A dataset has \(Q_1 = 30\), \(Q_3 = 50\).
85 is an outlier
(a) IQR \(= 20\); lower \(= 30 - 1.5 \times 20 = 0\); upper \(= 50 + 1.5 \times 20 = 80\).
(b) \(85 > 80\), so 85 is an outlier.
A dataset has \(Q_1 = 15\), \(Q_3 = 27\).
0 is not an outlier
(a) IQR \(= 12\); lower \(= 15 - 1.5 \times 12 = -3\); upper \(= 27 + 1.5 \times 12 = 45\).
(b) \(0 > -3\) (and \(< 45\)), so 0 is not an outlier.
Level 1 · Fluency
Which measure of centre is “resistant” to outliers?
The median.
Which measure of spread is resistant to outliers — the range or the IQR?
The IQR.
Is the mean or the median more affected by one extreme value?
The mean.
Level 2 · Application
Explain why an outlier changes the mean more than the median.
The mean uses the actual value of every data point, so an extreme value shifts the total and the mean; the median depends only on the position of the middle value, which barely moves.
Explain why the range is affected by an outlier but the IQR usually is not.
The range is max − min, so an extreme value changes it directly; the IQR uses only the quartiles (the middle of the data), which an outlier at the end does not reach.
Explain why adding one very large value raises the mean but leaves the mode unchanged.
The mean uses the total of all values, so a large value increases it; the mode is just the most frequent value, which a single new value does not change.
Level 3 · Further Application
Test marks are 60, 62, 65, 68, 70. A mark of 10 is added.
(a) Before: \(325 \div 5 = 65\). After: \(335 \div 6 = 55.83\) — the mean falls by about 9.
(b) The median changes only slightly (from 65 to 63.5), showing it is far less affected.
Scores are 40, 42, 45, 47, 51. A score of 5 is added.
(a) Before: \(225 \div 5 = 45\). After: \(230 \div 6 = 38.33\) — the mean falls by about 7.
(b) The median changes only slightly (from 45 to 43.5), showing it is far less affected.
Weights (kg) are 20, 21, 22, 24, 25. A weight of 60 is added.
(a) Before: \(112 \div 5 = 22.4\). After: \(172 \div 6 = 28.67\) — the mean rises by about 6.
(b) The median changes only slightly (from 22 to 23), so it is far less affected by the outlier.
Level 1 · Fluency
A headline says “Average income up 10%”. Why might “average” be misleading?
The mean can be inflated by a few very high incomes, so most people’s income may not have risen by 10%.
A report says “the average house price is $1.2 million”. Why might “average” overstate a typical price?
A few very expensive houses can inflate the mean, so the typical (median) price may be much lower.
An advertisement claims “up to 50% off”. Why might this be misleading?
“Up to” means the biggest discount possible — most items may have a much smaller discount, so it overstates the typical saving.
Level 2 · Application
A company advertises an “average” salary of $95 000, but most staff earn $60 000. Explain how this can happen.
A few executives on very high salaries pull the mean up; the median (typical) salary is much lower than the advertised mean.
A gym advertises that members “lose on average 8 kg”. Explain how this average could be misleading.
A few members with very large losses can pull the mean up, so most members may lose far less than 8 kg; the median would give a fairer typical figure.
A property report states the “average” rent rose 12%. Explain how a small number of properties could cause this.
A few very large rent increases can raise the mean substantially, even if most rents changed little — so the mean overstates the typical rise.
Level 3 · Further Application
An advertisement claims “our customers save on average $500”.
(a) Whether it is a mean or median, the spread of savings, and how the sample was chosen (a few large savers can inflate a mean).
(b) The median saving — it is not distorted by a small number of unusually large savings.
An advertisement claims “our students improve by an average of 20 marks”.
(a) Whether it is a mean or median, the spread of improvements, and how the students were selected (a few very large gains can inflate a mean).
(b) The median improvement — it is not distorted by a few unusually large gains.
A headline reads “average wait time is only 5 minutes”.
(a) Whether it is the mean or median, the spread of wait times, and when/where the data was collected (a few very short waits can lower a mean).
(b) The median wait — it better represents a typical wait because it is not pulled down by a few very short waits.
Level 1 · Fluency
Name the three features used to summarise a distribution.
Shape, centre and spread.
Name two measures that describe the centre of a distribution.
The mean and the median.
Is “IQR” a measure of centre or a measure of spread?
A measure of spread.
Level 2 · Application
Two datasets have the same median but Dataset A has a much larger IQR. Compare their spread.
Dataset A is more spread out (greater variability) around the centre, while Dataset B is more tightly clustered.
Two datasets have the same mean, but Dataset A has a much larger range than Dataset B. Compare their spread.
Dataset A is more spread out (greater variability), while Dataset B’s values are more tightly clustered around the mean.
Two datasets have the same spread, but Dataset A has a higher median than Dataset B. Compare their centre.
Dataset A is centred on higher values (its typical value is greater), while their variability is the same.
Level 3 · Further Application
Class P: median 65, IQR 10, symmetric. Class Q: median 65, IQR 22, positively skewed.
(a) Same centre (median 65), but Class Q is much more spread out (IQR 22 vs 10).
(b) A positive skew suggests a tail of higher marks — a few students scored well above the rest.
Group M: median 50, IQR 8, symmetric. Group N: median 50, IQR 20, negatively skewed.
(a) Same centre (median 50), but Group N is much more spread out (IQR 20 vs 8).
(b) A negative skew suggests a tail of lower values — a few members scored well below the rest.
Store A: median $40, IQR $6, symmetric. Store B: median $55, IQR $6, positively skewed.
(a) Same spread (IQR $6), but Store B has a higher centre (median $55 vs $40).
(b) A positive skew suggests a tail of higher values — a few unusually large amounts.
Level 1 · Fluency
A dataset has one clear peak. What is its modality?
Unimodal.
A dataset has two clear peaks. What is its modality?
Bimodal.
A dataset has three or more clear peaks. What is its modality?
Multimodal.
Level 2 · Application
A histogram of café visits shows two clear peaks — one at breakfast, one at lunch. State the modality and what it suggests.
Bimodal — it suggests two busy periods (breakfast and lunch rushes).
A histogram of gym attendance shows two peaks — one in the early morning and one in the evening. State the modality and what it suggests.
Bimodal — it suggests two busy periods (before and after work).
A dataset of exam marks shows a single clear peak around 65. State the modality and what it suggests.
Unimodal — it suggests most students scored near one central value (about 65).
Level 3 · Further Application
A dataset of ages at a family event has peaks near 10, 40 and 70.
(a) Multimodal (three peaks).
(b) It suggests three generations are present — children, parents and grandparents — each forming a cluster of ages.
A histogram of shoe sizes sold has peaks at children’s and adults’ sizes.
(a) Bimodal (two peaks).
(b) It suggests two distinct groups of buyers — children and adults — each clustering around their own typical sizes.
A histogram of race finishing times has peaks near 30, 45 and 60 minutes.
(a) Multimodal (three peaks).
(b) It suggests three groups of runners — for example fast, average and slower runners — each forming a cluster of times.
Level 1 · Fluency
In a positively skewed distribution, the longer tail is on which side?
The right (higher-value) side.
In a negatively skewed distribution, the longer tail is on which side?
The left (lower-value) side.
In a symmetric distribution, how do the mean and median compare?
They are approximately equal.
Level 2 · Application
For a positively skewed dataset, state whether the mean is greater than or less than the median, and why.
The mean is greater than the median — the high-value tail pulls the mean up above the middle value.
For a negatively skewed dataset, state whether the mean is greater than or less than the median, and why.
The mean is less than the median — the low-value tail pulls the mean down below the middle value.
A dataset has mean 50 and median 50. State the likely shape and why.
Approximately symmetric — the mean and median being equal indicates no skew in either direction.
Level 3 · Further Application
A dataset of house prices has mean $920k and median $780k.
(a) Positively (right) skewed.
(b) The mean ($920k) is well above the median ($780k), which happens when a tail of expensive houses pulls the mean upward.
A dataset of test marks has mean 54 and median 66.
(a) Negatively (left) skewed.
(b) The mean (54) is below the median (66), which happens when a tail of low marks pulls the mean downward.
A dataset of waiting times has mean 14 min and median 9 min.
(a) Positively (right) skewed.
(b) The mean (14 min) is above the median (9 min), which happens when a tail of long waits pulls the mean upward.
Level 1 · Fluency
For skewed data with outliers, which pair is usually reported — mean & SD, or median & IQR?
Median and IQR (both resist outliers).
For roughly symmetric data with no outliers, which pair is usually reported — mean and SD, or median and IQR?
Mean and standard deviation.
Which measure of spread should be reported alongside the median?
The interquartile range (IQR).
Level 2 · Application
A dataset is roughly symmetric with no outliers. State which measures of centre and spread are appropriate and why.
Mean and standard deviation — with symmetric, outlier-free data they use all the values and describe it well.
A dataset is strongly skewed with several outliers. State which measures of centre and spread are appropriate and why.
Median and IQR — they are resistant to outliers and skew, so they describe the data more fairly than the mean and standard deviation.
Explain why the mean and standard deviation suit symmetric, outlier-free data.
They use every value, so with symmetric, outlier-free data they summarise the centre and spread accurately without being distorted.
Level 3 · Further Application
A dataset of incomes is strongly right-skewed with a few very high values.
(a) Median and IQR.
(b) The IQR-based rule \((Q_3 + 1.5 \times IQR)\) is based on quartiles, which are not distorted by the extreme high incomes, so it flags outliers reliably.
A dataset of house prices is strongly right-skewed with a few very expensive homes.
(a) Median and IQR.
(b) The IQR rule \((Q_3 + 1.5 \times IQR)\) is based on quartiles, which the extreme high prices do not distort, so it flags outliers reliably.
A dataset of adult heights is roughly symmetric with no outliers.
(a) Mean and standard deviation — the data is symmetric and outlier-free, so these use all values and describe it well.
(b) If the data became skewed or gained outliers, switch to the median and IQR.
Level 1 · Fluency
Find the median of 3, 8, 9, 11, 15.
9
The middle value = 9.
Find the median of 4, 6, 10, 13, 18.
10
The middle value \(=\) 10.
Find the median of 2, 5, 7, 8, 9, 12.
7.5
Two middle values 7 and 8: \(\dfrac{7 + 8}{2} =\) 7.5.
Level 2 · Application
The skewed data 2, 3, 3, 4, 30 is given. State the most appropriate measure of centre and calculate it.
3
Median (data is skewed by 30). Ordered middle value = 3.
The skewed data 1, 2, 2, 3, 40 is given. State the most appropriate measure of centre and calculate it.
2
Median (the value 40 makes the data skewed). Ordered middle value \(=\) 2.
The skewed data 5, 6, 7, 8, 50 is given. State the most appropriate measure of centre and calculate it.
7
Median (50 is an outlier that would inflate the mean). Ordered middle value \(=\) 7.
Level 3 · Further Application
A dataset is 5, 6, 6, 7, 8, 9, 40.
(a) Mean \(= 81 \div 7 = 11.57\); median \(= 7\).
(b) The median (7) — the value 40 is an outlier that inflates the mean, so the median is more representative.
A dataset is 3, 4, 5, 5, 6, 7, 40.
mean 10; median 5
(a) Mean \(= 70 \div 7 = 10\); median \(= 5\).
(b) The median (5) — the value 40 is an outlier that inflates the mean, so the median is more representative.
A dataset is 8, 9, 10, 11, 12, 13, 90.
mean 21.86; median 11
(a) Mean \(= 153 \div 7 = 21.86\); median \(= 11\).
(b) The median (11) — the value 90 is an outlier that inflates the mean, so the median better represents the data.
Level 1 · Fluency
On a box-plot, the line inside the box represents which value?
The median (here, 8).
On a box-plot, the left and right edges of the box represent which two values?
The lower quartile \(Q_1\) (left) and the upper quartile \(Q_3\) (right).
On a box-plot, the ends of the two whiskers represent which values (assuming no outliers)?
The minimum and the maximum.
Level 2 · Application
From the box-plot above, state the five-number summary.
Minimum 2, \(Q_1\) 5, median 8, \(Q_3\) 12, maximum 20.
A box-plot has its box from 10 to 22 with the median line at 16, and whiskers reaching 4 and 30. State the five-number summary.
Minimum 4, \(Q_1\) 10, median 16, \(Q_3\) 22, maximum 30.
A box-plot has whiskers from 20 to 60, a box from 30 to 50, and the median at 42. State the IQR and the range.
IQR 20; range 40
IQR \(= Q_3 - Q_1 = 50 - 30 = 20\); range \(= 60 - 20 = 40\).
Level 3 · Further Application
Two groups’ results are shown as parallel box-plots.
(a) Group 1 median \(= 9\); Group 2 median \(= 12\) — Group 2 has the higher median.
(b) Group 1 IQR \(= 13 - 6 = 7\); Group 2 IQR \(= 15 - 9 = 6\) — Group 1 is slightly more spread out in the middle 50%.
Two classes’ marks give these five-number summaries. Class A: 4, 9, 13, 18, 25. Class B: 8, 14, 16, 20, 28.
Class B median higher; Class A IQR 9, Class B IQR 6
(a) Class A median \(= 13\), Class B median \(= 16\) — Class B has the higher median.
(b) Class A IQR \(= 18 - 9 = 9\); Class B IQR \(= 20 - 14 = 6\) — Class A’s middle 50% is slightly more spread out.
Two teams’ scores give these five-number summaries. Team X: 2, 6, 10, 14, 18. Team Y: 5, 8, 11, 13, 20.
Team Y median slightly higher; ranges almost equal
(a) Team X median \(= 10\), Team Y median \(= 11\) — Team Y’s median is slightly higher.
(b) Team X range \(= 18 - 2 = 16\); Team Y range \(= 20 - 5 = 15\) — almost the same, with Team X marginally more spread out.
Level 1 · Fluency
What five values make up a five-number summary?
Minimum, lower quartile (\(Q_1\)), median, upper quartile (\(Q_3\)), maximum.
In a five-number summary, which value sits between \(Q_1\) and \(Q_3\)?
The median (\(Q_2\)).
Which two values of a five-number summary are used to find the range?
The minimum and the maximum (range \(=\) max \(-\) min).
Level 2 · Application
For the data 3, 5, 7, 8, 9, 12, 15, 18, 20, write the five-number summary.
Min 3; \(Q_1 = 6\); median \(= 9\); \(Q_3 = 16.5\); max 20.
For the data 2, 4, 6, 7, 10, 13, 16, write the five-number summary.
Min 2; \(Q_1 = 4\); median \(= 7\); \(Q_3 = 13\); max 16.
For the data 5, 8, 11, 14, 15, 19, 22, 26, write the five-number summary.
Min 5; \(Q_1 = 9.5\); median \(= 14.5\); \(Q_3 = 20.5\); max 26.
Level 3 · Further Application
A dataset has minimum 4, maximum 29, median 15, \(Q_1\) 9 and \(Q_3\) 21.
not
(a) Range \(= 29 - 4 = 25\); IQR \(= 21 - 9 = 12\).
(b) Upper boundary \(= 21 + 1.5 \times 12 = 39\); \(29 < 39\), so 29 is not an outlier.
A dataset has minimum 6, maximum 48, median 24, \(Q_1\) 16 and \(Q_3\) 34.
not
(a) Range \(= 48 - 6 = 42\); IQR \(= 34 - 16 = 18\).
(b) Upper boundary \(= 34 + 1.5 \times 18 = 61\); \(48 < 61\), so 48 is not an outlier.
A dataset has minimum 2, maximum 60, median 20, \(Q_1\) 14 and \(Q_3\) 30.
outlier
(a) Range \(= 60 - 2 = 58\); IQR \(= 30 - 14 = 16\).
(b) Upper boundary \(= 30 + 1.5 \times 16 = 54\); \(60 > 54\), so 60 is an outlier.
Level 1 · Fluency
Two datasets have the same range but different IQRs. What does the IQR tell you that the range does not?
The IQR describes the spread of the middle 50% of the data, ignoring extreme values, whereas the range depends only on the two extremes.
Two datasets have the same median but different ranges. What does the range tell you here?
The range shows how far apart the extreme (largest and smallest) values are, i.e. the total spread of each dataset.
If one dataset has an outlier and another does not, which summary statistic is most affected?
The mean (and the range); the median and IQR are resistant.
Level 2 · Application
Dataset A: median 50, IQR 12. Dataset B: median 50, IQR 25. Compare their consistency.
Both have the same centre, but Dataset A is more consistent (its middle 50% is packed into a smaller IQR of 12 vs 25).
Dataset A: median 30, range 20. Dataset B: median 30, range 45. Compare their spread.
Both have the same centre, but Dataset B is far more spread out — its range (45) is more than double Dataset A’s (20).
Dataset A: median 60, IQR 8, one outlier. Dataset B: median 60, IQR 8, no outliers. Compare their reliability.
They have the same centre and middle spread, but Dataset B is more reliable because it has no outlier distorting it.
Level 3 · Further Application
Team X: median 20, IQR 6, range 15, no outliers. Team Y: median 22, IQR 14, range 40, one outlier.
(a) Team Y has a slightly higher median (22 vs 20) but a much larger spread (IQR 14 vs 6; range 40 vs 15) and an outlier.
(b) Team X is more reliable — its smaller IQR and range and absence of outliers show more consistent results.
Machine A: median 500, IQR 5, range 12, no outliers. Machine B: median 502, IQR 15, range 40, one outlier.
(a) Machine B has a slightly higher median (502 vs 500) but a much larger spread (IQR 15 vs 5; range 40 vs 12) and an outlier.
(b) Machine A is more reliable — its smaller IQR and range and absence of outliers show more consistent output.
Class P: median 70, IQR 8, range 24, no outliers. Class Q: median 68, IQR 20, range 55, one outlier.
(a) Class P has a slightly higher median (70 vs 68) and a much smaller spread (IQR 8 vs 20; range 24 vs 55) with no outliers.
(b) Class P performed more consistently — its smaller IQR and range and lack of outliers show less variation.
Level 1 · Fluency
If one box-plot sits entirely to the right of another, what does that indicate?
That dataset has generally higher values (a higher centre and higher extremes).
If two box-plots have the same median but one has a much wider box, what does the wider box show?
That dataset has a larger IQR — its middle 50% of values is more spread out.
If one box-plot’s whiskers are much longer than another’s, what does that indicate?
That dataset has a larger range — its values are more spread out overall.
Level 2 · Application
Two towns’ rainfall box-plots are shown; Town A’s box is higher and narrower than Town B’s. Interpret this in context.
Town A has higher typical rainfall (higher median) and more consistent rainfall (narrower box/IQR); Town B is lower and more variable.
Two athletes’ race-time box-plots are shown; Athlete A’s box is lower and narrower than Athlete B’s. Interpret this in context.
Athlete A has faster typical times (lower median) and more consistent times (narrower box/IQR); Athlete B is slower and more variable.
Two shops’ daily-sales box-plots are shown; Shop A’s median is higher but its box is much wider than Shop B’s. Interpret this in context.
Shop A has higher typical sales (higher median) but less consistent sales (wider box/larger IQR); Shop B sells less but more steadily.
Level 3 · Further Application
Parallel box-plots compare exam marks for a morning class and an afternoon class.
(a) The second class performed better — its median (12) and quartiles are higher than the first class’s (median 9).
(b) The second class was slightly more consistent (IQR 6 vs 7), so its middle 50% of marks were a little less spread out.
Two classes’ marks give these five-number summaries. Morning: 5, 9, 12, 15, 20. Afternoon: 8, 13, 16, 19, 24.
Afternoon better; equal consistency
(a) The afternoon class performed better — its median (16) and quartiles are higher than the morning class’s (median 12).
(b) Both had the same IQR (\(15 - 9 = 6\) and \(19 - 13 = 6\)), so their middle 50% of marks were equally consistent.
Two branches’ wait times (min) give these five-number summaries. Branch A: 2, 5, 8, 11, 15. Branch B: 3, 4, 6, 8, 10.
Branch B faster and more consistent
(a) Branch B is faster — its median wait (6 min) is lower than Branch A’s (8 min), and its whole box sits at lower times.
(b) Branch B is more consistent — its IQR (\(8 - 4 = 4\) min) is smaller than Branch A’s (\(11 - 5 = 6\) min).