Skip to content

Analysing: Evaluating Data and Conclusions

Syllabus mapping

Working Scientifically outcome — Analysing (part 2 of 2). Covers distinguishing accuracy/precision/reliability/validity, identifying genuine limitations in a method, suggesting targeted improvements, and drawing a conclusion that answers the original hypothesis. See Trends and Relationships for the first half of Analysing.

Exact NESA outcome code TODO — confirm against the official syllabus PDF (Resources) before treating any wording on this page as verbatim NESA text.

You need to know: four words — accuracy, precision, reliability, validity — that sound interchangeable in everyday speech and are absolutely not interchangeable in a Stage 6 evaluation, plus how to write a conclusion and an improvement that actually says something specific.

Core

Accuracy, precision, reliability, validity

These four are the single most commonly confused set of terms in Working Scientifically, and "Evaluate" questions frequently hinge on picking the right one.

Term What it actually means
Accuracy How close a measured/calculated value is to the true or accepted value.
Precision How close repeated measurements are to each other — regardless of whether they're close to the true value.
Reliability Whether repeating the method (by you, or by someone else) would produce consistent results.
Validity Whether the method actually tests what it claims to test — controlled variables genuinely held constant, no confounding factor sneaking in.
Worked example — telling them apart

A set of repeated timing trials for a pendulum's period all cluster tightly around 2.30 s (precise), but the accepted period for that length is actually 1.99 s — perhaps the student was timing every second oscillation without realising it (inaccurate, and specifically because of a systematic error). The method, if repeated exactly as performed, would keep giving the same wrong answer (still reliable — reliability doesn't require being right, just being consistent). If the pendulum's swing amplitude was accidentally allowed to vary between trials, the method itself would also have a validity problem, independent of the timing mistake.

It's entirely possible to be precise, reliable, and invalid or inaccurate all at once — which is exactly why an evaluation needs to check all four separately, not treat "the results were consistent" as proof the experiment was good.

Identifying genuine limitations

A limitation names a specific weakness in the method and explains how it affected the result — not a vague, generic phrase.

Weak (avoid) Strong (specific)
"Human error" "Reaction time when starting/stopping the stopwatch by hand introduces a timing uncertainty of roughly ±0.2 s per reading, which is a large fraction of a ~2 s period."
"Equipment wasn't accurate enough" "The ruler's 1 mm resolution limited the precision of each length measurement to ±0.5 mm, which was the dominant source of uncertainty in the final gradient."
"More trials would help" "Only three repeats were taken at each length, giving a relatively wide half-range uncertainty; five or more repeats would narrow the uncertainty in the mean without requiring different equipment."

The pattern: name the specific source → explain the mechanism → connect it to the actual effect on the result. "Human error" could mean a dozen different things and explains nothing on its own.

Suggesting improvements that match the limitation

An improvement should fix the specific limitation just identified, not restate "use better equipment" as a generic close-out line. If the limitation was reaction-time uncertainty in hand-timing, the improvement is a light gate or photogate timer — not "be more careful." If the limitation was too narrow a tested range to confirm a trend, the improvement is testing a wider range of the independent variable — not repeating the same range more times.

Drawing a conclusion

A conclusion should explicitly answer the hypothesis stated back in Questioning and Predicting — supported, not supported, or partially supported — and say why, referencing the actual result.

Worked example — a complete conclusion

"The hypothesis — that increasing the incline angle increases the trolley's acceleration, because the component of gravity parallel to the slope increases with angle — was supported. Acceleration increased from \(1.7 \pm 0.2\ \text{m s}^{-2}\) at 10° to \(6.1 \pm 0.3\ \text{m s}^{-2}\) at 50°, and a graph of acceleration against \(\sin\theta\) was linear through the origin (see Trends and Relationships), consistent with the predicted \(g\sin\theta\) relationship. The gradient of that graph, \(9.6 \pm 0.4\ \text{m s}^{-2}\), agrees with the accepted value of \(g\) within experimental uncertainty."

Notice what makes this a strong conclusion rather than a restatement of the hypothesis: it cites actual numbers, references the graph's shape as evidence (not just "the results matched"), and explicitly compares the extracted gradient against the accepted value — tying together every earlier stage of the investigation, from Questioning and Predicting through to here.

Advanced

A result that doesn't support the hypothesis is not a failed evaluation. If a hypothesis isn't supported, the strongest response identifies why, distinguishing between "the underlying physics is different from what I assumed" and "a limitation in the method masked the real relationship." Concluding "the hypothesis was not supported, most likely due to [specific limitation], rather than because the underlying relationship doesn't hold" is a genuinely sophisticated Band 6-level evaluation — it shows you can separate the physics from the execution of the experiment.

The two boxes below are optional, tertiary-level statistical tools for a more rigorous evaluation — genuinely beyond this syllabus, but exactly the kind of thing the strongest students (often also doing Extension 1/2 Maths) reach for in a depth study. Pick whichever is relevant. See also the standard error and t-test boxes on Trends and Relationships, which these build on.

Extension — Standard deviation and the 68–95–99.7 rule

Beyond the Physics 11–12 syllabus — won't appear in the HSC, included for interest / depth study inspiration.

📎 Depth study idea

Standard deviation isn't just a number to quote — for data whose random error is roughly symmetric (a reasonable assumption for most Year 11 measurement error), it lets you judge how unusual a particular reading is, using the empirical rule:

  • About 68% of readings fall within \(\pm 1\sigma\) of the mean
  • About 95% fall within \(\pm 2\sigma\)
  • About 99.7% fall within \(\pm 3\sigma\)

This turns "is this reading an anomaly?" (see Conducting) into a checkable statement rather than a gut feeling.

Worked example

From the pendulum data in Trends and Relationships (see its "Standard error of the mean" Extension box): mean \(\bar{T} = 1.992\) s, \(\sigma \approx 0.030\) s. A sixth trial reads \(2.15\) s — how unusual is that?

\[ \frac{2.15 - 1.992}{0.030} \approx 5.3 \]

That reading sits about \(5.3\sigma\) from the mean — far outside the \(\pm 3\sigma\) range that should contain 99.7% of genuine random variation. That's strong statistical grounds to treat it as an anomaly worth investigating (see Conducting), not just an unusually large but legitimate result.

Extension — Confidence intervals: how sure, precisely?

Beyond the Physics 11–12 syllabus — won't appear in the HSC, included for interest / depth study inspiration.

📎 Depth study idea

"Report a mean and an uncertainty" and "state how confident you are that the true value lies in a given range" are closely related but not identical. A confidence interval makes the second one explicit, combining the standard error with a critical t-value (see Trends and Relationships) to give a range you can attach an actual confidence level to:

\[ \bar{x} \pm t_{crit} \times SE \]
Worked example

Using the pendulum data again: \(\bar{T} = 1.992\) s, \(SE = 0.014\) s, \(df = 4\), \(t_{crit}(95\%) = 2.776\).

\[ 1.992 \pm (2.776 \times 0.014) = 1.992 \pm 0.039\ \text{s} \]

So: "We are 95% confident the true period lies between 1.953 s and 2.031 s." That's a substantially more precise and defensible claim than just writing \(\pm\) some uncertainty — it states exactly what the range means and how sure you are of it, which is usually what "evaluate the reliability of your result" is really asking for.

Extension — Type A and Type B uncertainty

Beyond the Physics 11–12 syllabus — won't appear in the HSC, included for interest / depth study inspiration.

📎 Depth study idea

Professional research reports this same idea using the language of Type A and Type B uncertainty evaluation (ISO/BIPM's Guide to the Expression of Uncertainty in Measurement): Type A uncertainty comes from the statistical spread of repeated measurements (what Precision and Uncertainty mostly deals with), while Type B comes from everything else — manufacturer specifications, calibration certificates, prior knowledge, judgement. A depth study evaluation that explicitly separates its sources of uncertainty into Type A and Type B, rather than lumping "all the reasons this might be wrong" into one list, is a genuinely university-level treatment of exactly this section's content.

Video/visual resources

  • 🎥 Khan Academy — TODO: source an accuracy vs precision explainer. Should ideally also cover reliability and validity in the same video (all four terms together, since they're most often confused as a group) — a dartboard-style visual analogy is the standard, effective way this gets taught. Essential.
  • 🎥 Physics High — TODO: check, but this is a general science-skills topic rather than physics-specific — a general science channel may have better dedicated coverage than a physics one.

Check yourself

  1. A group's five repeated measurements of a spring's extension are 4.5 cm, 8.9 cm, 4.6 cm, 4.4 cm, 4.5 cm. Explain what the outlier at 8.9 cm suggests about the reliability of their measuring technique on that particular trial, and what it does not by itself tell you about accuracy.

    Answer

    The one wildly different value suggests that trial specifically was not measured reliably/consistently with the others (something went wrong on that repeat — likely worth flagging as an anomaly, see Conducting) — the other four are tightly clustered and can be treated as reliable. It doesn't, by itself, say anything about accuracy: even the tightly clustered value of ~4.5 cm could still be inaccurate if there's an unnoticed systematic error (e.g. a mis-zeroed ruler).

  2. Rewrite this limitation to make it specific: "There was human error in reading the values, which affected the results."

    Answer

    E.g. "Reading the ruler at an angle rather than directly perpendicular to the scale (parallax error) could have introduced a small, direction-dependent inaccuracy in each length measurement, of up to a few millimetres depending on the viewing angle used." — names the specific mechanism (parallax), and estimates its likely size/effect rather than leaving it vague.

  3. A student's conclusion reads: "My results proved my hypothesis was correct." Identify two things wrong with this as a scientific conclusion.

    Answer

    (1) Experiments support or are consistent with a hypothesis — they don't "prove" it; a hypothesis remains open to being disproven by future evidence. (2) The conclusion doesn't reference any actual data, graph, or comparison to a predicted/accepted value — a strong conclusion has to show how the results support the hypothesis, not just assert that they do.

  4. (Stretch — uses the Extension boxes above.) A data set has mean \(50.0\) and \(\sigma = 2.0\). Using the 68–95–99.7 rule, state the range within which about 95% of readings should fall, and explain whether a new reading of \(53.5\) would be unusual.

    Answer

    95% of readings should fall within \(\pm 2\sigma\), i.e. between \(46.0\) and \(54.0\). A reading of \(53.5\) sits within that range (it's \(1.75\sigma\) from the mean), so it would not be considered statistically unusual — it's consistent with normal random variation, not grounds to flag it as an anomaly.