Year 12 · Statistical analysis
Quick tips — memory joggers
Scatterplots & association
Correlation coefficient \(r\)
Least-squares line \(y = a + bx\)
Predicting & causation
Level 1 · Fluency
On a scatterplot investigating how rainfall affects umbrella sales, which variable goes on the horizontal axis?
The independent variable (rainfall).
You are plotting how the number of hours studied affects a test score. Which variable goes on the horizontal (\(x\)) axis?
Hours studied (the independent variable).
The independent variable is plotted on the horizontal axis, so hours studied goes on the \(x\)-axis and the test score on the \(y\)-axis.
When constructing a scatterplot from paired data, should the plotted points be joined with a line?
No.
Each pair is plotted as a single point and the points are not joined; a line of best fit may be added afterwards to show the trend.
Level 2 · Application
Given paired data \((1, 6)\), \((2, 8)\), \((3, 11)\), \((4, 12)\), describe how you would plot it as a scatterplot.
Plot each pair as a point, with the first value on the \(x\)-axis and the second on the \(y\)-axis; do not join the points.
A table gives paired data \((2, 5)\), \((4, 9)\), \((6, 10)\), \((8, 14)\). Describe how you would plot it as a scatterplot.
Plot each pair as a point using the first value as \(x\) and the second as \(y\); leave the points unjoined.
Mark points at \((2, 5)\), \((4, 9)\), \((6, 10)\) and \((8, 14)\), scaling both axes to fit the data. The rising points suggest a positive association, but the points are not joined.
A table lists five students’ hours of sleep and their reaction times. Explain how you would turn the table into a scatterplot.
Take sleep as \(x\) and reaction time as \(y\); plot each student as one point \((\text{sleep}, \text{reaction})\) and do not join them.
Sleep is the independent variable so it goes on the horizontal axis; reaction time is the dependent variable on the vertical axis. Plot one point per student and do not join the points.
Level 3 · Further Application
A scatterplot of umbrellas sold against daily rainfall is shown.
(a) A strong, positive, roughly linear association — as rainfall increases, umbrella sales increase.
(b) No — scatterplot points are individual observations and should not be joined; a line of best fit is drawn instead.
A scatterplot of weekly ice-cream sales against average temperature shows points rising steadily from lower-left to upper-right in a tight band.
(a) A strong, positive, linear association.
(b) No — the points are separate observations and should not be joined.
(a) The points rise together (positive), lie along a line (linear) and cluster in a tight band (strong).
(b) Joining the points would wrongly imply values between observations were measured; instead a line of best fit is used to show the trend.
A scatterplot of the fuel used by a car against the distance travelled shows points lying very close to a rising straight line.
(a) A strong, positive, linear association.
(b) Yes — a straight line of best fit is appropriate.
(a) More distance uses more fuel (positive), the points follow a line (linear) and lie very close to it (strong).
(b) Because the pattern is clearly linear, a straight line of best fit models it well.
Level 1 · Fluency
A scatterplot’s points rise from lower-left to upper-right. Is the association positive or negative?
Positive.
A scatterplot’s points fall from upper-left to lower-right. Is the association positive or negative?
Negative.
A scatterplot’s points are spread out with no upward or downward trend. What form of association is this?
No (linear) association.
Level 2 · Application
Describe, with justification, the association in a scatterplot whose points fall steadily from upper-left to lower-right in a tight band.
A strong, negative, linear association — as \(x\) increases \(y\) decreases, and the tight band shows the points lie close to a straight line.
Describe, with justification, the association in a scatterplot whose points rise from lower-left to upper-right but are widely scattered.
A weak, positive, linear association — as \(x\) increases \(y\) tends to increase (positive), roughly along a line (linear), but the wide scatter shows the points do not cluster closely, so it is weak.
Describe, with justification, the association in a scatterplot whose points rise and lie very close to a straight line.
A strong, positive, linear association — the points rise (positive) along a line (linear) and cluster tightly about it (strong).
Level 3 · Further Application
A scatterplot shows points scattered with no clear upward or downward trend.
(a) No (or very weak) association — the points do not follow an upward or downward line.
(b) One variable would be a poor predictor of the other, since there is no clear relationship.
A scatterplot of daily temperature against hot-chocolate sales shows points falling from upper-left to lower-right in a fairly tight band.
(a) A strong, negative, linear association — as temperature rises, sales fall (negative), along a line (linear), and the tight band shows it is strong.
(b) Temperature would be a good predictor of hot-chocolate sales, because the strong association means points lie close to the line.
A scatterplot shows points that rise along a clear curve that gets steeper — not a straight line.
(a) A positive but non-linear (curved) association — \(y\) increases with \(x\), but along a curve rather than a straight line.
(b) No — a straight line of best fit would not model a curved trend well.
Level 1 · Fluency
An association where points lie almost exactly on a rising straight line is described as which strength?
Strong (and positive, linear).
Points loosely scattered around a falling line — is the strength strong, moderate or weak?
Weak (and negative, linear).
Points lie on a clear curve rather than a straight line. Is the form linear or non-linear?
Non-linear.
Level 2 · Application
Describe an association using the three descriptors (direction, form, strength) for points that loosely trend downward along a line.
Negative direction, linear form, weak-to-moderate strength.
Describe an association using the three descriptors for points that rise steeply and cluster tightly along a line.
Positive direction, linear form, strong strength.
Points form a clear U-shape (falling then rising). Describe the association using the three descriptors as far as they apply.
The form is non-linear (curved), so a single direction and strength do not apply — it is a non-linear association with no consistent positive or negative direction.
Level 3 · Further Application
Two studies report correlations: Study A points form a tight rising line; Study B points rise but are widely scattered.
(a) A: positive, linear, strong. B: positive, linear, weak.
(b) Study A — its stronger association means predictions from the line are more reliable.
Two datasets are compared: dataset P points fall tightly along a line; dataset Q points fall but are widely scattered.
(a) P: negative, linear, strong. Q: negative, linear, weak.
(b) Dataset P — its stronger (tighter) association makes predictions from the line more reliable.
Two datasets are compared: dataset X shows points with no clear trend; dataset Y points lie on a tight rising line.
(a) X: no (or very weak) linear association. Y: positive, linear, strong.
(b) Dataset Y — a line of best fit only suits a clear linear pattern, which X does not have.
Level 1 · Fluency
In “temperature affects ice-cream sales”, which is the dependent variable?
Ice-cream sales (they depend on temperature).
In “rainfall affects crop yield”, which is the independent variable?
Rainfall (the yield depends on it).
In “the more hours worked, the higher the pay”, which is the dependent variable?
Pay (it depends on the hours worked).
Level 2 · Application
A study looks at how study hours affect exam marks. Identify the independent and dependent variables.
Independent: study hours; dependent: exam marks.
A study examines how a car’s speed affects its stopping distance. Identify the independent and dependent variables.
Independent: speed; dependent: stopping distance.
A study looks at how the number of daily sunlight hours affects a plant’s growth. Identify the independent and dependent variables.
Independent: sunlight hours; dependent: plant growth.
Level 3 · Further Application
A researcher investigates whether a town’s height above sea level affects its average temperature.
(a) Independent: height above sea level; dependent: average temperature.
(b) Height on the \(x\)-axis, temperature on the \(y\)-axis.
An economist studies whether a country’s average income affects its life expectancy.
(a) Independent: average income; dependent: life expectancy.
(b) Income on the \(x\)-axis, life expectancy on the \(y\)-axis.
A gym studies whether a runner’s weekly training hours affect their 5 km race time.
(a) Independent: weekly training hours; dependent: 5 km race time.
(b) Training hours on the \(x\)-axis, race time on the \(y\)-axis.
Level 1 · Fluency
Give an example of two variables likely to have a positive association.
e.g. a person’s height and their shoe size (taller people tend to have larger feet).
Give an example of two variables likely to have a negative association.
e.g. a car’s age and its value (older cars are generally worth less).
Give an example of two variables likely to have little or no association.
e.g. a person’s shoe size and their exam mark.
Level 2 · Application
Give an example of two real-world variables likely to have a negative association, and explain why.
e.g. a car’s age and its value — older cars are generally worth less, so as age increases value decreases.
Give an example of two real-world variables likely to have a positive association, and explain why.
e.g. hours studied and exam mark — more study time tends to raise the mark, so as one increases the other increases.
A news report says “suburbs with more parks tend to have higher house prices”. State the two variables and the likely direction of association.
Number of parks and house price; a positive association (more parks, higher prices).
Level 3 · Further Application
A news article claims “more screen time is linked to lower sleep”.
(a) Screen time and hours of sleep; a negative association is claimed.
(b) Other factors (e.g. lifestyle, homework) could affect both — correlation does not establish causation.
A report claims “students who eat breakfast score higher in exams”.
(a) Whether a student eats breakfast and their exam score; a positive association.
(b) A third factor (e.g. a stable home routine or more sleep) may cause both, so correlation is not causation.
An article states “suburbs with more fast-food outlets have higher obesity rates”.
(a) Number of fast-food outlets and obesity rate; a positive association.
(b) A lurking variable such as average income or lifestyle may drive both, so correlation does not prove causation.
Level 1 · Fluency
What is the value of Pearson’s \(r\) for a perfect negative linear relationship?
\(r = -1\).
What is the value of Pearson’s \(r\) for a perfect positive linear relationship?
\(r = +1\).
What value of Pearson’s \(r\) indicates no linear association?
\(r = 0\).
Level 2 · Application
A dataset gives \(r = 0.86\). Describe the strength and direction of the linear association.
A strong, positive linear association (\(r\) is close to \(+1\)).
A dataset gives \(r = -0.42\). Describe the strength and direction of the linear association.
A weak, negative linear association (\(|r| < 0.5\), and the sign is negative).
A dataset gives \(r = -0.93\). Describe the strength and direction of the linear association.
A strong, negative linear association (\(r\) is close to \(-1\)).
Level 3 · Further Application
Two predictors of a region’s temperature are compared: distance from the coast gives \(r = -0.55\); altitude gives \(r = -0.90\).
(a) Distance from coast: moderate negative association; altitude: strong negative association.
(b) Altitude — its \(r\) (\(-0.90\)) is closer to \(-1\), indicating a stronger linear relationship and more reliable predictions.
Two predictors of a student’s exam mark are compared: hours studied gives \(r = 0.78\); hours of TV gives \(r = -0.30\).
(a) Hours studied: strong positive association; hours of TV: weak negative association.
(b) Hours studied — its \(|r| = 0.78\) is larger (closer to 1), so it predicts the mark more reliably.
For a group of adults, height vs arm span gives \(r = 0.95\), while height vs age gives \(r = 0.10\).
(a) Height vs arm span: strong positive association; height vs age: very weak (almost no) linear association.
(b) Height vs arm span — \(r = 0.95\) is much closer to \(1\) than \(0.10\).
Level 1 · Fluency
A line of best fit summarises the trend of a scatterplot. Should it pass through roughly the middle of the points?
Yes — it should balance the points, with about as many above as below.
A line of best fit slopes downward from left to right. Is the association positive or negative?
Negative.
Roughly how should the data points be arranged about a good line of best fit?
About equal numbers above and below the line, close to it.
Level 2 · Application
A line of best fit is \(y = 3x + 5\). Describe what the gradient tells you about the association.
For each 1-unit increase in \(x\), \(y\) increases by about 3 — a positive association.
A line of best fit is \(y = -2x + 40\). Describe what the gradient tells you about the association.
For each 1-unit increase in \(x\), \(y\) decreases by about 2 — a negative association.
A line of best fit is \(y = 0.5x + 10\). Describe what the gradient tells you about the association.
For each 1-unit increase in \(x\), \(y\) increases by about 0.5 — a positive (but gentle) association.
Level 3 · Further Application
A line of best fit for umbrellas sold (\(y\)) against daily rainfall (\(x\) mm) is \(y = 2.5x + 3\).
18 umbrellas
(a) Positive — each extra mm of rainfall is associated with about 2.5 more umbrellas sold.
(b) \(y = 2.5 \times 6 + 3 =\) 18 umbrellas.
A line of best fit for a plant’s height (\(y\) cm) against the number of weeks (\(x\)) is \(y = 4x + 2\).
22 cm
(a) Positive — the plant grows about 4 cm for each extra week.
(b) \(y = 4 \times 5 + 2 = 20 + 2 =\) 22 cm.
A line of best fit for a phone’s battery (\(y\) %) against hours of use (\(x\)) is \(y = -12x + 100\).
64%
(a) Negative — the battery falls about 12% for each extra hour of use.
(b) \(y = -12 \times 3 + 100 = -36 + 100 =\) 64%.
Level 1 · Fluency
Which method gives a single, exact line of best fit — fitting by eye or least-squares by technology?
Least-squares by technology (fitting by eye can differ between people).
True or false: fitting a line by eye always gives exactly the same line for everyone.
False — different people place a by-eye line slightly differently.
Which method finds the line that minimises the total squared vertical distances to the points?
The least-squares method (using technology).
Level 2 · Application
Explain one advantage of using technology rather than the eye to fit a line of best fit.
Technology gives a precise, reproducible least-squares line that minimises the prediction errors, rather than a subjective estimate.
Explain one advantage of first fitting a line of best fit by eye before using technology.
A by-eye line gives a quick check that the trend is roughly linear and that the technology’s line is reasonable (a sensible gradient and intercept).
Explain why two people fitting a line of best fit by eye may get different gradients.
Fitting by eye is a subjective judgement, so each person balances the points slightly differently, producing different gradients.
Level 3 · Further Application
Two students fit a line to the same scatterplot by eye and get slightly different lines.
(a) Fitting by eye is a judgement, so people place the line slightly differently.
(b) The least-squares method (via technology) computes one definite line that minimises the total squared vertical distance to the points, giving a consistent result.
A class fits a line by eye to the same data and the gradients range from 1.8 to 2.6.
(a) By-eye fitting is subjective, so each student positions the line differently.
(b) Least-squares technology computes one definite line that minimises the total squared vertical distance to the points, giving the same result for everyone.
A student’s by-eye line is \(y = 3x + 1\); technology gives \(y = 2.7x + 1.4\).
(a) The by-eye line is only an estimate, so it differs from the calculated line.
(b) The technology (least-squares) line \(y = 2.7x + 1.4\) should be reported, because it is precise, reproducible and minimises the prediction errors.
Level 1 · Fluency
A least-squares regression line has equation \(y = a + bx\). What does \(b\) represent?
The gradient (slope) of the line.
In the least-squares line \(y = a + bx\), what does \(a\) represent?
The \(y\)-intercept — the value of \(y\) when \(x = 0\).
In \(y = a + bx\), if \(b\) is negative, does \(y\) increase or decrease as \(x\) increases?
\(y\) decreases.
Level 2 · Application
Technology gives the least-squares line \(y = 32 - 0.009x\) for temperature \(y\) against altitude \(x\) (metres). Use it to predict the temperature at 400 m.
28.4°C
\(y = 32 - 0.009 \times 400 = 32 - 3.6 =\) 28.4°C.
Technology gives the least-squares line \(y = 5 + 1.4x\) for a plant’s height \(y\) (cm) against the week \(x\). Predict the height in week 8.
16.2 cm
\(y = 5 + 1.4 \times 8 = 5 + 11.2 =\) 16.2 cm.
A least-squares line \(y = 120 - 4x\) models a phone’s battery \(y\) (%) against hours used \(x\). Predict the battery level after 7 hours.
92%
\(y = 120 - 4 \times 7 = 120 - 28 =\) 92%.
Level 3 · Further Application
A least-squares line for ice-cream sales (\(y\)) against temperature (\(x\)°C) is \(y = -15 + 8x\).
161 sales
(a) \(y = -15 + 8 \times 22 = -15 + 176 =\) 161 sales.
(b) Each 1°C rise in temperature is associated with about 8 more sales.
A least-squares line for weekly sales \(y\) (in $1000s) against advertising spend \(x\) (in $1000s) is \(y = 18 + 2.5x\).
$33 000 (i.e. \(y = 33\))
(a) \(y = 18 + 2.5 \times 6 = 18 + 15 = 33\), i.e. $33 000 in sales.
(b) Each extra $1000 spent on advertising is associated with about $2500 more in sales.
A least-squares line \(y = 15 + 6x\) models the number of chirps \(y\) a cricket makes per minute against the temperature \(x\) (°C).
165 chirps/min
(a) \(y = 15 + 6 \times 25 = 15 + 150 =\) 165 chirps/min.
(b) Each 1°C rise in temperature is associated with about 6 more chirps per minute.
Level 1 · Fluency
In \(y = 50 + 8x\) (cost $\(y\) for \(x\) items), what does the intercept 50 represent?
The fixed cost when \(x = 0\) (before any items) — $50.
In \(y = 20 + 3x\) (taxi fare $\(y\) for \(x\) km), what does the gradient 3 represent?
The cost per kilometre — the fare rises by $3 for each extra km.
In \(y = 100 - 5x\) (water \(y\) litres in a tank after \(x\) minutes), what does the gradient \(-5\) mean?
The tank loses 5 litres of water each minute.
Level 2 · Application
For \(y = 29.2 - 0.011x\) (temperature vs height in metres), interpret the gradient in context.
For every 1 metre increase in height above sea level, the average temperature falls by about 0.011°C.
For \(y = 15 + 0.8x\) (phone bill $\(y\) for \(x\) minutes of calls), interpret the intercept 15 in context.
The fixed monthly charge of $15 that applies even when no calls are made (\(x = 0\)).
For \(y = 500 - 20x\) (value $\(y\) of a machine after \(x\) years), interpret the gradient in context.
The machine loses about $20 in value for each additional year.
Level 3 · Further Application
A regression line for weekly wage $\(y\) against years of experience \(x\) is \(y = 900 + 45x\).
(a) Each additional year of experience is associated with about $45 more weekly wage.
(b) The intercept $900 is the predicted wage at 0 years’ experience (a starting wage); it is meaningful here since \(x = 0\) is realistic.
A regression line for a taxi fare $\(y\) against distance \(x\) (km) is \(y = 4 + 2.2x\).
(a) Each extra kilometre travelled adds about $2.20 to the fare.
(b) The intercept $4 is the flag-fall (base) charge at 0 km; it is meaningful here since \(x = 0\) represents starting the trip.
A regression line for a baby’s mass \(y\) (kg) against age \(x\) (months) is \(y = 3.4 + 0.6x\).
(a) The baby gains about 0.6 kg for each additional month of age.
(b) The intercept 3.4 kg is the predicted mass at 0 months, i.e. the birth mass; it is meaningful here since \(x = 0\) is realistic.
Level 1 · Fluency
Predicting within the range of the data is called interpolation or extrapolation?
Interpolation.
Predicting outside the range of the data is called interpolation or extrapolation?
Extrapolation.
A line is fitted to data with \(x\) from 5 to 20. Is predicting at \(x = 25\) interpolation or extrapolation?
Extrapolation (\(x = 25\) is beyond the range 5–20).
Level 2 · Application
A line \(y = 2.5x + 3\) is fitted to data with \(x\) from 1 to 8. Predict \(y\) at \(x = 5\) and state whether it is interpolation or extrapolation.
\(y = 2.5 \times 5 + 3 = 15.5\). Since \(x = 5\) is inside 1–8, this is interpolation.
A line \(y = 4x - 2\) is fitted to data with \(x\) from 2 to 10. Predict \(y\) at \(x = 7\) and state whether it is interpolation or extrapolation.
\(y = 26\); interpolation
\(y = 4 \times 7 - 2 = 28 - 2 =\) 26. Since \(x = 7\) is inside 2–10, this is interpolation.
A line \(y = 100 - 6x\) is fitted to data with \(x\) from 0 to 12. Predict \(y\) at \(x = 9\) and state which type of prediction this is.
\(y = 46\); interpolation
\(y = 100 - 6 \times 9 = 100 - 54 =\) 46. Since \(x = 9\) is inside 0–12, this is interpolation.
Level 3 · Further Application
The line \(y = 2.5x + 3\) is fitted to data with \(x\) from 1 to 8.
33
(a) \(y = 2.5 \times 12 + 3 =\) 33.
(b) It is extrapolation (\(x = 12\) is beyond the data range 1–8), so it is less reliable — the linear trend may not continue outside the observed data.
The line \(y = 8 + 3x\) is fitted to data with \(x\) from 1 to 9.
53
(a) \(y = 8 + 3 \times 15 = 8 + 45 =\) 53.
(b) It is extrapolation (\(x = 15\) is beyond the data range 1–9), so it is less reliable — the linear trend may not continue past the observed data.
The line \(y = 200 - 4x\) models a tank’s water (litres) over \(x\) minutes, and was fitted for \(x\) from 0 to 30.
20 litres
(a) \(y = 200 - 4 \times 45 = 200 - 180 =\) 20 litres.
(b) It is extrapolation (\(x = 45\) is beyond the data range 0–30), so it is unreliable — the tank may already be empty, so the linear model need not hold.
Level 1 · Fluency
Which is generally more reliable — interpolation or extrapolation?
Interpolation (predicting inside the data range).
Is extrapolation predicting inside or outside the range of the data?
Outside the range of the data.
Why is interpolation usually reliable?
Because it predicts within the range where the linear relationship was observed to hold.
Level 2 · Application
Explain why extrapolating a line of best fit far beyond the data can give a misleading prediction.
The linear relationship is only known to hold within the data collected; beyond that range the true relationship may change, so the prediction may be far from reality.
Explain one limitation of interpolation, even though it is usually reliable.
Interpolation still assumes the linear trend holds exactly; if the association is weak (points scattered), even an interpolated value can be inaccurate.
A line of best fit predicts a negative number of sales for a very high price. Explain what this shows about the model’s limits.
The linear model breaks down outside sensible values — a negative number of sales is impossible, showing that extrapolating too far can give meaningless results.
Level 3 · Further Application
A model of plant height \(y\) (cm) against week \(x\) is reliable for weeks 1–10.
(a) Week 30 is well outside the data (1–10), so the model may not apply and could over-predict.
(b) The plant eventually stops growing (reaches a maximum height), which a straight line cannot show.
A model of a runner’s speed \(y\) against training weeks \(x\) is reliable for weeks 1–12.
(a) Week 40 is far beyond the data (1–12), so the linear trend may no longer hold and the prediction could be too high.
(b) The runner’s improvement eventually levels off (plateaus), which a straight line cannot show.
A model of temperature \(y\) against altitude \(x\) is reliable up to 2000 m.
(a) 8000 m is far outside the data range (up to 2000 m), so the relationship may change and the prediction is unreliable.
(b) At high altitudes other conditions (such as air pressure) change, so the temperature need not keep falling at the same steady rate.
Level 1 · Fluency
If taller people tend to weigh more, what type of association is this?
A positive association.
As a car gets older its value tends to fall. What type of association is this?
A negative association.
Shoe size and exam mark show no clear pattern. What type of association is this?
No (or negligible) association.
Level 2 · Application
A regression line \(y = 0.9x + 12\) relates weight \(y\) (kg) to height above 150 cm (\(x\)). Predict the weight for a height 20 cm above 150 cm.
30 kg
\(y = 0.9 \times 20 + 12 =\) 30 kg (predicted, for that model).
A regression line \(y = 1.2x + 30\) relates a plant’s height \(y\) (cm) to the number of weeks \(x\). Predict the height at 15 weeks.
48 cm
\(y = 1.2 \times 15 + 30 = 18 + 30 =\) 48 cm.
A regression line \(y = 2x + 5\) relates points scored \(y\) to minutes played \(x\). Predict the points for 25 minutes played.
55 points
\(y = 2 \times 25 + 5 = 50 + 5 =\) 55 points.
Level 3 · Further Application
Data on a car’s age \(x\) (years) and value $\(y\) gives the line \(y = 32\,000 - 2500x\), with \(r = -0.95\).
(a) \(r = -0.95\) shows a strong negative linear association — value falls steadily as age increases.
(b) \(y = 32\,000 - 2500 \times 6 =\) $17 000; the strong correlation makes this prediction reasonably reliable within the data range.
Data on a house’s distance \(x\) (km) from the city and its price $\(y\) (in $1000s) gives the line \(y = 900 - 25x\), with \(r = -0.88\).
$650 000
(a) \(r = -0.88\) shows a strong negative linear association — price falls as distance from the city increases.
(b) \(y = 900 - 25 \times 10 = 900 - 250 = 650\), i.e. $650 000; the strong correlation makes this reasonably reliable within the data range.
Data on a runner’s weekly training \(x\) (hours) and their 5 km time \(y\) (minutes) gives the line \(y = 40 - 1.5x\), with \(r = -0.82\).
28 minutes
(a) \(r = -0.82\) shows a strong negative linear association — more training is linked to a faster (smaller) time.
(b) \(y = 40 - 1.5 \times 8 = 40 - 12 =\) 28 minutes; the strong correlation makes this reasonably reliable within the data range.
Level 1 · Fluency
Before fitting a line, what should you check the scatterplot for?
Whether the association is roughly linear (a line of best fit only suits linear patterns).
After plotting the data, which numerical measure quantifies the strength of a linear association?
Pearson’s correlation coefficient \(r\).
If a scatterplot shows a clear curve, is a straight line of best fit suitable?
No — a straight line does not suit a non-linear (curved) pattern.
Level 2 · Application
A scatterplot of sales vs advertising spend shows a clear rising linear trend. State the next step to model it and what you would report.
Fit a least-squares regression line, then report its equation, the direction/strength of the association and \(r\).
A scatterplot of monthly heating cost vs average temperature shows a clear falling linear trend. State the next step to model it and what you would report.
Fit a least-squares regression line, then report its equation, the negative direction/strength of the association and \(r\).
A scatterplot of test score vs hours of sleep shows points that curve (rising then falling). Explain why a straight line of best fit is unsuitable and what you should report instead.
The trend is non-linear, so a straight line would fit poorly; report that the association is non-linear rather than fitting and using a single straight line.
Level 3 · Further Application
For the umbrella scatterplot above (umbrellas sold vs rainfall):
10.5 ≈ 11 umbrellas
(a) A strong, positive, linear association — more rainfall, more umbrellas sold.
(b) \(y = 2.5 \times 3 + 3 =\) 10.5 ≈ 11 umbrellas.
A study records average temperature \(x\) (°C) and cold-drink sales \(y\). The scatterplot rises steeply in a tight band, and the fitted line is \(y = 5x + 20\).
170 drinks
(a) A strong, positive, linear association — warmer days, more cold drinks (points in a tight band, so strong).
(b) \(y = 5 \times 30 + 20 = 150 + 20 =\) 170 drinks.
A study records distance \(x\) (km) from a city and average house price \(y\) (in $1000s). The scatterplot falls in a tight band, and the fitted line is \(y = 800 - 30x\).
$440 000
(a) A strong, negative, linear association — further from the city, lower price (tight band, so strong).
(b) \(y = 800 - 30 \times 12 = 800 - 360 = 440\), i.e. $440 000.
Level 1 · Fluency
Give one reason personal data should be kept private in a study.
To protect individuals’ identities and comply with privacy/consent requirements.
Give one reason a survey sample should represent the whole population.
So the results are not biased and can be generalised to the whole population.
What should researchers obtain from people before collecting their personal data?
Informed consent.
Level 2 · Application
A health study collects data from volunteers at one clinic. Explain one source of bias.
Volunteers from a single clinic may not represent the whole population (self-selection and location bias), so results may not generalise.
An online poll about internet use is only answered by frequent internet users. Explain the bias this creates.
Self-selection (coverage) bias — heavy internet users are over-represented and light or non-users are missed, so the results overstate internet use and do not represent everyone.
A survey on cultural practices is written only in English. Explain one issue with its cultural responsiveness.
People who do not read English are excluded, which biases the sample and fails to respect and include all cultural groups.
Level 3 · Further Application
A company wants to collect and publish customer location data to study buying patterns.
(a) Obtaining informed consent, and protecting privacy (e.g. not identifying individuals or revealing sensitive locations).
(b) Aggregate and anonymise the data so no individual can be identified before analysing or publishing it.
A school wants to publish a study linking students’ home postcodes to their marks.
(a) Individual students could be identified from a postcode combined with marks, and consent is needed to use their results.
(b) Group postcodes into broad regions and anonymise the data before publishing so no student can be identified.
A fitness app sells its users’ health and location data to advertisers.
(a) Using the data beyond what users consented to, and the risk of identifying or tracking individuals from health and location data.
(b) Only share aggregated, anonymised data, and obtain clear user consent before sharing anything.
Level 1 · Fluency
Name one source of reliable published data for a bivariate investigation.
A government statistics agency (e.g. the Australian Bureau of Statistics) or another official public dataset.
Name one pair of biometric measurements you could investigate for an association.
e.g. a person’s height and arm span (or hand span and foot length).
After collecting two numerical variables, what plot would you draw to look for an association?
A scatterplot.
Level 2 · Application
You download data on rainfall and crop yield across regions. Describe how you would check for an association.
Plot yield against rainfall on a scatterplot, describe the direction/form/strength, and calculate \(r\) to quantify the linear association.
You collect students’ height and arm span. Describe how you would check for an association.
Plot arm span against height on a scatterplot, describe the direction/form/strength, and calculate \(r\) to quantify the linear association.
You download government data on the unemployment rate and average income by region. Describe how you would check for an association.
Plot average income against the unemployment rate on a scatterplot, describe the direction/form/strength, and calculate \(r\) to quantify the linear association.
Level 3 · Further Application
A student uses public data on income and life expectancy across countries.
(a) Likely positive; life expectancy is the dependent variable, income the independent variable.
(b) Other factors (e.g. healthcare, diet, education) may drive both — correlation is not causation.
A student uses biometric data on adults’ foot length and height.
(a) Likely positive; height is the dependent variable, foot length the independent variable.
(b) Both are driven by overall body growth (a common cause), so the correlation does not show one causes the other.
A student uses government data on a region’s average temperature and its electricity use for heating.
(a) Likely negative (colder regions use more heating); heating use is the dependent variable, temperature the independent variable.
(b) Other factors (e.g. income, house size, insulation) also affect heating use, so correlation is not causation.