The Inference Primerinductive inference · information · models

Case study

A case study: long-range precipitation forecasting

This page describes one historical application in enough detail to see the whole method at work: a state-space, a search for patterns, a set of conditional inferences, and a trial against data the model had not seen. It is presented as a methodological example, not as a position on any question of climate policy.

Stylised contour map of ocean temperature bands beside a mountain watershed profile

The setting

At the end of the 1970s, skill in forecasting precipitation faded within weeks. A study reported by Ronald Christensen and colleagues in 1980, carried out under contract for water-resources research with California as the area of interest, asked whether the entropy-minimax approach to pattern discovery could extend useful skill to a year or more. The available record was roughly 126 years of observations, which had to serve both for building the model and for testing it.

That constraint is the first thing worth noticing. A century of annual observations is a small sample by any modern standard — of the order of a hundred cases, from which both the patterns and the honest estimate of their worth must come. The design decisions that follow are all shaped by that scarcity.

The state-spaces

The unobserved space was deliberately coarse: a year was classed wet if precipitation exceeded the multi-year median and dry otherwise. A two-state space is a modest question, and modest questions are the ones a hundred cases can answer. Attempting eight precipitation bands would have left a handful of cases per cell and produced patterns supported by nothing.

The observed space was assembled from quantities available before the forecast was due: Pacific sea-surface temperatures in specified regions and seasons, prior-year precipitation at named stations, and tree-ring growth as a proxy for earlier conditions. The lags mattered as much as the variables — each observation had to be genuinely in hand at forecast time, which is the discipline described under observed and unobserved spaces.

The result, in the form it takes

The reported model consisted of three patterns, each a conjunction of conditions on those observed variables, each carrying a conditional inference of the form given this pattern, the probability that next year is wet is p ± e. The patterns were defined so as to be mutually exclusive: the second applied only where the first did not, and the third only where neither did. Together they partitioned the observed space, so every year fell under exactly one rule.

The reported probabilities departed appreciably from the roughly even base rate implied by a median split — some well below, one above — with stated uncertainties wide enough to be honest about the sample sizes behind them. The study reported statistically significant forecasts in an independent validation trial at lead times of the order of one to three years, with the significance strongest in extreme years.

Those figures are given here as what the source reports. This primer has not re-run the analysis, and the reader should treat the numbers as the literature's claim rather than as an independently verified fact. The original contract reports, and a later peer-reviewed paper on seasonal forecasting by the same information-theoretic route in Monthly Weather Review (1985), are listed in the bibliography.

Why the validation trial is the interesting part

With three rules over a century of data, the search space of candidate conjunctions was large relative to the sample. As explained under pattern discovery, that guarantees some apparent patterns by chance alone, and it is why the design reserved data for a trial in which the finished rules were applied to years that played no part in forming them. Whatever weight the result carries, it carries because of that trial, not because the rules fit the years they were derived from.

Two features of the study deserve credit on this score: the outcome definition was fixed in advance by a median rule rather than chosen to flatter the model, and the rules were reported in full, so that anyone could apply them to later years. Publishing a model in a form that can be run forward by a third party is a strong test to volunteer for.

The El Nino connection

One durable by-product was the prominence the study gave to Pacific sea-surface temperature — that is, to the oscillation now known as El Nino–Southern Oscillation — as a predictor of North American precipitation at long lead times. That connection is now thoroughly established through many independent lines of work, and it underpins operational seasonal outlooks today; NOAA's explainer on ENSO and the Physical Sciences Laboratory's ENSO resources set out the current understanding and data.

It is fair to say that the study was working in that direction early. It is not fair to credit it with the discovery, and the wider record — with independent contributions from many groups through the 1980s and after — should not be collapsed into any single study's account.

How to read a result like this

  • Look at the question before the score. Skill at predicting a median split is a weaker claim than the same skill at predicting an amount, and both may be reported as "forecasting precipitation".
  • Ask what was withheld, and when. The value of a trial depends entirely on the withholding having preceded the search.
  • Ask how many rules were considered. Three reported rules may be the survivors of thousands examined; the selection is part of the method and belongs in the report.
  • Read the interval. A stated uncertainty of ±0.10 on a probability estimate is telling you how few cases stand behind it.

Applied to this study, those questions produce a measured verdict: a carefully designed, honestly validated piece of work on a very small sample, whose method is more instructive than its numbers, and whose central variable has been vindicated many times over by later research done independently of it.