Missing information
This page does the one piece of derivation the primer really needs. It takes the additivity of measures, applies it to Shannon's measure, and arrives at a quantity that can reasonably be called the information an inference is missing.
Setting up
Let Sh(·) denote Shannon's measure of a set. Let X be an observed state-space and Y an unobserved one, each treated as a set. Then:
- Y ∩ X is what the two spaces share — what observing X settles about Y.
- Y − X is what remains of Y once the shared part is removed — what observing X leaves open.
The identity
Additivity gives, for the measure of a set difference:
Sh(Y − X) = Sh(Y) − Sh(Y ∩ X)
Each term has a standard name. Sh(Y) is the entropy of the unobserved space: everything there was to settle before any observation. Sh(Y ∩ X) is the information that the observed space carries about the unobserved one — mutual information, in the usual vocabulary. And Sh(Y − X) is the conditional entropy: what is left.
The identity is a bookkeeping statement, and its force comes from the reading. Since Sh(Y) is fixed once the question is fixed, the conditional entropy varies inversely with the information: every bit the observation supplies is a bit no longer outstanding.
Why the residue is an inference
The step worth pausing on is the identification of the set difference Y − X with the inference itself. The argument runs like this. An inference is the move from what is observed to what is not; what makes the move non-trivial is exactly the part of Y that X does not settle; that part is Y − X; and the functional form of Sh(Y − X) is the standard conditional-entropy function. So the inference has a measure, and the measure is the conditional entropy.
Treat this as an interpretation supported by a derivation rather than as a theorem about inference. It is a good interpretation — it gets the limiting cases exactly right, which is more than most proposals in this area manage — but the reader should know which parts are mathematics and which are reading.
The two endpoints
The missing information ranges between zero and Sh(Y).
- Zero. The observation settles the question completely; the inference has become a deduction. This is the sense in which deduction is the limiting case of induction rather than a separate faculty.
- Maximum. X and Y share nothing, so Y − X reduces to Y and the conditional entropy reduces to the entropy. The observation was irrelevant and the inference asserts nothing beyond what was known before it.
Between the two, every inference has a location, and that is what makes comparison possible: two candidate models, applied to the same unobserved space, can be ranked by how much they leave missing.
What "missing" is missing for
The phrase is the information missing in this inference, for a deductive conclusion, and the last clause is essential. The quantity is not a general measure of ignorance; it is the shortfall relative to the specific standard of certainty about Y. Change the unobserved space and the shortfall changes, because the standard has changed. A model that leaves little missing about a coarse question may leave a great deal missing about a fine one, and neither number alone tells you whether the model is good — only the pair, question and shortfall, does.
Where this leads
Having a measure of an inference makes optimisation possible: choose the model that minimises what is left missing, subject to honesty about what the data support. That is the programme examined under the principles of reasoning, and the reason it needed this page first is that no such programme can even be stated until the thing being optimised has been defined. The derivation of the underlying measure is Shannon's; the use made of it here belongs to the inductive-inference literature that followed.