Outcome
The outcome of an event is the state that obtains in the unobserved state-space — an elemental state in the space formed by a model's dependent variables, or an abstraction from several such states. In the standing example, rain is the outcome when the space is {rain, no rain}.
A separate word for a separate thing
Ordinary usage lets "prediction" stand for both the claim and the event. Keeping them apart is worth a page of its own, because almost every failure of model validation is a failure to keep them apart.
The inference is what the model asserts before the event: a distribution over the unobserved space, conditioned on what was observed. The outcome is what the world then does. The two live at different times, are known by different people, and — crucially — the outcome must be recorded without reference to the inference. When the recording of outcomes is influenced by what was predicted, the model is grading its own work.
Ways this goes wrong in practice
- The outcome is defined after the fact. If the boundary between wet year and dry year, or between failure and degradation, is settled once the data are in view, the definition can drift toward whatever makes the model look right. Fixing the unobserved space before the trial is not bureaucracy; it is the whole of the control.
- The outcome is partly caused by the prediction. A maintenance forecast that triggers an intervention changes the thing it predicted. A risk score that reallocates clinical attention does the same. The model is then being scored on a world it altered, and the honest response is to model the intervention too, not to ignore it.
- The outcome is not observed for every case. Cases lost to follow-up, tanks never dug up, rods never inspected — these are missing outcomes, and they are rarely missing at random. The subset with recorded outcomes can support a model that fails badly on the rest.
- The outcome is coarsened to make the model look better. Collapsing eight rainfall bands into two after seeing the results is a real and common manoeuvre. It is legitimate only if decided in advance, because coarsening always raises apparent accuracy and always lowers what the model actually says.
Outcomes and the dependent variables
A model's dependent variables span the unobserved space, and the outcome is a state in that space. When several dependent variables are in play, the space is their joint product, and an outcome may be described either at the elementary level or as an abstraction over several elementary states — any failure mode rather than this particular failure mode. Both are legitimate; they answer different questions, and a model built for one should not be quietly scored on the other.
Why the outcome is the only unforgiving test
Every internal measure of a model can, with enough freedom, be made to look good. Fit improves with parameters. Cross-validated fit improves with enough attempts at cross-validation. What cannot be argued with is a run of outcomes recorded after the model was frozen, on cases the model has never seen. That is the discipline described under pattern discovery, and it is why the case study in this primer is presented in terms of its validation trial rather than its fit. The wider methodological literature on scoring rules and calibration is surveyed accessibly in the philosophy of statistics literature.