The Inference Primerinductive inference · information · models

Method

The problem of induction

Every predictive model asserts something the evidence does not entail. The question of what licenses that assertion is the problem of induction, and it is not a philosopher's curiosity: it is the thing a modeller decides, by act if not by argument, every time a model is chosen.

A sequence of identical marks continuing toward a horizon where the next mark is uncertain

The problem, stated plainly

Suppose every observed instance of some kind has had a property. What licenses the claim that the next instance will have it too? Not deduction: no contradiction follows from supposing the next instance different. Not experience, without circularity: to argue that induction has worked before and so will work again is itself an induction, and so presupposes what it sets out to establish.

David Hume put the difficulty in essentially this form in the eighteenth century, and it has not been dissolved since. The literature is enormous; the Stanford Encyclopedia's survey is the best free map of it.

Why it is a practical problem

The philosophical version can look like a puzzle to be admired and set aside. The working version cannot, because it appears in a concrete form: many different models fit the data equally well and disagree about the next case. The record does not choose among them. Something else must, and whatever does the choosing is the modeller's answer to Hume, whether stated or not.

Usually it is not stated. It is embedded in defaults: this family of curves rather than that, this many parameters, this prior, this penalty, these variables because they were in the database. Each of those is an inductive commitment. Their being invisible does not make them innocent; it makes them unexaminable, which is worse.

The standard responses

Simplicity

Prefer the simplest model consistent with the evidence. Ancient, widely used, and — as a bare principle — underdetermined, because simplicity depends on the language in which the model is expressed. A curve that is simple in one coordinate system is baroque in another. Made precise via description length, the idea becomes genuinely usable, and it then shows its kinship with the information measures in this primer: the shortest description and the least missing information are close relatives.

Probabilistic coherence

Represent belief as a probability distribution and update by conditioning. This is the dominant contemporary response, and it delivers a great deal: coherence constraints, a calculus for combining evidence, convergence results under repeated observation. What it does not deliver is the starting point. The prior is an inductive commitment; the framework tells you how to move, not where to stand. Bayesian epistemology is the entry point to that literature.

Falsificationism

Deny that induction is needed: conjecture boldly, test severely, retain what survives. This has been enormously productive as a norm of scientific conduct, and it is thin as an account of prediction. Sooner or later somebody must act on a theory that has merely survived, and the step from "not yet refuted" to "act on it" is inductive again.

Explicit optimisation of an information measure

Fix a measure of what an inference leaves undetermined, then let a stated optimisation principle choose the model. This is the tradition these pages describe. Its honest claim is not that induction has been justified but that the commitment has been relocated: from a modeller's undocumented taste to a written principle that can be criticised, compared with alternatives, and tested against outcomes. The stronger claim sometimes made — that the principles of reasoning have thereby been discovered, closing the problem — goes beyond what the wider literature accepts, and this primer does not make it.

What a good answer looks like in practice

Because no answer is available that removes risk, the working standard has to be procedural. Three properties are worth insisting on:

  1. Explicit. The commitment is written down as a principle, not left in the defaults of a software package.
  2. Uniform. The same principle is applied to every candidate model, so that the comparison measures the models rather than the modeller's enthusiasm.
  3. Testable. The finished model is scored on cases that played no part in building it, which is the discipline described under pattern discovery.

None of this is a solution. All of it is the difference between a model whose inductive step can be argued about and one whose inductive step is invisible. The next page takes up a specific proposal of that kind.