Uncertainty as a Tool
Probability gives scientific uncertainty a structure that can be examined. It helps us specify possible outcomes, relate evidence to competing explanations, and compare actions without pretending that incomplete information has disappeared.
The earlier article, Prediction Has Limits, examined the assumptions and measurements behind a model. This article develops the next question: what must accompany a probability estimate before it can responsibly inform a decision?
Related reading: Prediction Has Limits
A number needs a question
Start with an event: a defined outcome that either occurs or does not occur in the stated setting. A probability model assigns probabilities to events according to consistent mathematical rules. Those rules alone do not establish that the model describes the physical situation.[1]
For a robot, “success” might mean holding a specified object for three seconds without dropping it. For an educational activity, it might mean correctly explaining a particular concept on a specified assessment. Neither definition automatically covers a broader claim such as safe operation or lasting understanding.
Record the population or operating conditions, the forecast horizon, and the information available when the forecast was made. These choices determine what observations can meaningfully challenge it.
Conditioning changes the answer
Bayes’ rule expresses how a probability changes when specified evidence is taken into account. The probability of a hypothesis after observing evidence depends on its prior probability and on how likely that evidence would be under the hypothesis and its alternatives.[2]
Consider an invented equipment-screening example with 1,000 components. Suppose 100 are defective; a detector flags 80 of those and also flags 90 of the 900 sound components. Then 80 of 170 flagged components are defective: approximately 47.1%.
Detecting 80% of defective components therefore does not mean that 80% of flagged components are defective. The denominators differ. This arithmetic illustration assumes the stated counts and describes no actual detector.
A useful habit follows: whenever someone presents a percentage, ask which cases form its denominator and whether the desired question reverses the conditioning.
Give uncertainty a source
Measurement uncertainty concerns the values reasonably attributable to a measured quantity, given available information. It is distinct from a known error.[3] A reported measurement should identify the quantity, unit, procedure, and meaning of its uncertainty.
For a hypothetical length reported as 0.100 m with standard uncertainty 0.002 m, the uncertainty has the same unit as the measured quantity. Their ratio, 0.02 or 2%, is a dimensionless relative standard uncertainty. It does not by itself specify a 95% interval.
A probability estimate can also be uncertain because the dataset is small, conditions differ, or the model omits relevant structure. These possibilities call for different responses. Additional repetitions may improve an estimate of variability while leaving an unrecognized systematic offset untouched.
Naming the source of uncertainty helps a team choose between collecting more observations, improving a measurement, changing a model, or restricting the claim.
Evaluate the probabilities
A reliability diagram compares forecast probabilities with observed event proportions in groups of predictions. The diagonal represents agreement; systematic departures motivate investigation.[4]
Consider 100 invented attempts, each forecast to have a 70% success probability, with 50 successes observed. Under an independent binomial model, the 95% Wilson interval for the success proportion is approximately 40.4%–59.6%. This interval reflects sampling uncertainty; it does not include every possible measurement or model error.[5]
Calibration does not measure every aspect of usefulness. A constant forecast can match an overall event rate while failing to distinguish individual cases. Measures such as the Brier score assess probability prediction quality but do not isolate calibration alone.[6]
Keep training, calibration fitting, and final evaluation appropriately separated. If records belong to the same device or sequence, choose a split that reflects the generalization being claimed. Record which evidence influenced each development decision.
Interpret statistical evidence
A p-value is the probability, under a specified null model and its assumptions, of a test statistic at least as extreme as the observed one. It is not the probability that the hypothesis is true. Nor does statistical significance establish a practically important effect.[7]
A frequentist confidence interval describes a procedure’s repeated-sampling coverage under assumptions. A Bayesian credible interval describes posterior probability under a specified prior and model. A prediction interval addresses a future observation. These answer different questions and should be labeled accordingly.
Bayesian analysis does not remove the need to check whether a model generates plausible data patterns.[8] Frequentist methods likewise require attention to design and assumptions. Scientific judgment includes examining alternatives, reporting uncertainty, and distinguishing statistical evidence from a causal claim.
Make the next test useful
A practical review can begin with five questions:
- What exact event is being predicted, and over what horizon?
- What information and assumptions produced the probability?
- Which observations were reserved to evaluate it?
- Which conditions could change the relationship?
- What action follows, and what are the consequences of error?
For a proposed robot pilot, these questions turn a confidence display into a testable record. For a learning activity, they help students explain a forecast and a revision. For a sponsor, they make the intended evidence and limits of a project easier to assess.
Probability forecasts can deteriorate under changed data conditions, as demonstrated in classification benchmarks.[9] Evaluation must therefore remain connected to use. A model’s earlier performance supports a bounded claim, and a changed setting can require new evidence.
The constructive response to uncertainty is to identify an achievable next test. Scientific confidence grows when assumptions become visible, observations become informative, and conclusions remain proportionate to what has been learned.
About the author
Dr. Albert Tan Lie Sing is a Mathematical Physicist, AI Robotics Systems Architect, and STEAM Education Innovator. Through HERO Science and Technology and alberttls.us, he connects scientific reasoning with intelligent systems, research-driven education, and collaborative technology development.
For sponsors
Support the development of openly documented probability lessons and bounded evaluation exercises through HERO Science and Technology. A proposed sponsorship would fund teaching materials, reproducible examples, and transparent assessment criteria. Read the article to see how these deliverables make scientific reasoning visible; their educational impact would be evaluated rather than assumed.
References
[1] Massachusetts Institute of Technology. (2013). Lecture 2: Conditioning and Bayes’ rule. In Probabilistic systems analysis and applied probability. MIT OpenCourseWare. Source.
[2] Massachusetts Institute of Technology. (2013). Lecture 2: Conditioning and Bayes’ rule. In Probabilistic systems analysis and applied probability. MIT OpenCourseWare. Source.
[3] Joint Committee for Guides in Metrology. (2008). Evaluation of measurement data—Guide to the expression of uncertainty in measurement (JCGM 100:2008). Source.
[4] Scikit-learn developers. (n.d.). Probability calibration. Scikit-learn documentation. Retrieved September 14, 2026, from Source.
[5] National Institute of Standards and Technology. (n.d.). 7.2.4.1. Confidence intervals. In NIST/SEMATECH e-Handbook of statistical methods. Retrieved September 14, 2026, from Source.
[6] Scikit-learn developers. (n.d.). Probability calibration. Scikit-learn documentation. Retrieved September 14, 2026, from Source.
[7] American Statistical Association. (2016, March 7). American Statistical Association releases statement on statistical significance and p-values. Source.
[8] Gelman, A., & Shalizi, C. R. (2013). Philosophy and the practice of Bayesian statistics. British Journal of Mathematical and Statistical Psychology, 66(1), 8–38. Source.
[9] Ovadia, Y., et al. (2019). Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift. Advances in Neural Information Processing Systems, 32. Source.
Comments