Form of Life a public ethics instrument
Browse

Ethics system

Construct

epistemic calibration

This entry is a construct rather than a person: epistemic calibration is the property of a judge whose stated confidence matches their observed accuracy — of everything said with eighty per cent confidence, about eighty per cent turns out true. It is measurable, it is trainable, and it is distinct from being right: a well-calibrated judge can be frequently wrong so long as they said so in advance. The construct's practical edge is that it converts intellectual honesty from a disposition into a scored record.

Why this reference appears

Each of these is interpretive context. None of them creates a fact about this life, or settles motive, diagnosis, identity, recurrence, or moral success.

Claim philosophical lineage

  1. Machine opacity is part of the reading

    Calibration is what an opaque reader still owes. If the process cannot be inspected, the one remaining checkable property is whether its stated confidence tracks its accuracy over many outputs — a well-calibrated black box is auditable at the level of its claims even where it is not at the level of its reasoning. The claim about opacity is livable because this construct supplies the fallback obligation.

Focused framework lineage

  1. The Calibrated Lens That Questions Its Own Opacity

    The framework names itself calibrated, so this construct is the standard it invites. Calibration is a measured property of a long run of confidence-tagged claims, not a disposition or a tone — which means the review's job is to ask whether the lens produces the record that would let anyone check it, and this entry supplies what such a record would have to contain.

Through-line philosophical lineage 2

  1. Frame-correction over self-defence

    Defending a challenged claim raises the cost of having been wrong, and the calibration literature shows that raising that cost is what produces overconfidence in the first place. Correcting the frame is the alternative that keeps a claim revisable — it moves the exchange off the ground where a stake in being right distorts the estimate.

  2. Architecture protects relationships, rather than perception management

    Calibration is not achieved by trying to be humble; it is achieved by structures — recording predictions before outcomes, scoring them, and reviewing the score. The through-line's preference for architecture over impression is this construct's own finding: the forecasters who improve are the ones inside a system that keeps their record, not the ones who resolved to be careful.

Ideas, works, and debates

Works

A scoring rule, a long forecasting programme, and the bias literature behind both.

  • Brier's 1950 scoring rule supplies the measurement: a proper score that a forecaster minimizes only by reporting their true belief, which is what makes calibration checkable rather than declarative.
  • Tetlock's Expert Political Judgment (2005) and Superforecasting (2015) supply the empirical programme — expert confidence largely uncorrelated with accuracy, and a trainable set of habits that improves both.
  • Murphy's decomposition of the Brier score into calibration, resolution, and uncertainty supplies the distinction the popular usage loses: being calibrated is not the same as being informative.
  • The overconfidence literature, from Lichtenstein and Fischhoff onward, supplies the default the construct is written against: people's confidence intervals are systematically too narrow.

Central ideas

Four ideas, and the second is the one that keeps the first from being cheap.

  • Calibration: stated confidence matches observed frequency, measured across many claims rather than assessed on any one.
  • Resolution: the willingness to depart from the base rate. A forecaster who says fifty per cent about everything is perfectly calibrated and useless, which is why calibration alone is not the goal.
  • Proper scoring: a rule under which honest reporting is the score-minimizing strategy, so the incentive to hedge or to bluff is removed by the mathematics rather than by character.
  • Trainability: calibration improves with feedback, decomposition of questions, and the habit of recording estimates before outcomes are known.

Distinctive vocabulary

Five terms, and the first two are constantly conflated.

  • Calibration: confidence matching frequency. Not accuracy, and not caution.
  • Resolution: informativeness — how far judgments depart from the base rate. The property calibration alone does not supply.
  • Brier score: a proper scoring rule for probabilistic forecasts. Lower is better.
  • Overconfidence: intervals too narrow or probabilities too extreme relative to accuracy. Not arrogance.
  • Base rate: the unconditional frequency. The thing a low-resolution forecaster never leaves.

Debates and disagreements

The construct's disputes are about where it applies.

  • Whether it transfers beyond forecastable questions: calibration needs resolvable outcomes, and most interesting claims about a life, a relationship, or a value are not resolvable in that sense — which bounds the construct far more tightly than its enthusiasts allow.
  • Whether training generalizes: the forecasting improvements are real within tournaments and their transfer to ordinary judgment is asserted more than shown.
  • Whether calibration can be gamed: a judge who selects easy questions scores well without being a better judge, which is why resolution is part of the decomposition.
  • Whether it privileges the quantifiable: attaching numbers to beliefs is a discipline for a specific class of claims, and treating it as the general form of honesty smuggles in a scope restriction.

Intellectual relationships

It is the measurable member of a family this record documents.

  • Fallibilism, documented here, is the disposition; calibration is its scored version — the difference between holding beliefs revisably and being able to show that one does.
  • Falsification-minded calibration, documented separately here, is the sibling construct: this one measures confidence against outcomes, that one organizes inquiry around what would disconfirm.
  • Media opacity, also here, is why calibration matters for machine readings — it is the property that remains checkable when the process is not.
  • Moral psychology supplies the pressure: overconfidence and motivated reasoning are the defaults this construct's machinery is designed to detect.

How it changes this reading

Four placements, all about what a claim owes when it cannot be inspected.

  • For machine opacity, it supplies the fallback obligation — an uninspectable reader can still be held to whether its confidence tracks its accuracy.
  • For the calibrated-lens framework's review, it supplies the standard the framework's own name invites, and the reminder that calibration is a record rather than a manner.
  • For frame-correction and architecture, it supplies the mechanism: calibration comes from systems that keep the record, not from resolving to be careful.
  • The guard, and it is a real limit here: most of this record's claims are not resolvable propositions, so the construct's measurement cannot be applied to them. It is cited for the discipline it names and not as a scoring of anything in this corpus.

Useful comparisons

Against neighbouring virtues.

  • Against accuracy: being right against knowing how likely one is to be right — a judge can have either without the other.
  • Against humility: calibration is a measured property, and a humble judge who is systematically underconfident is miscalibrated in the other direction.
  • Against precision: narrower estimates are better only if they remain calibrated, which is exactly what overconfidence violates.

Where the ideas meet

What the literature agrees on.

  • Confidence is measurable: it is not a mood but a claim with a checkable track record.
  • The default is overconfident: across domains and expertise levels, intervals are too narrow.
  • Feedback is the mechanism: improvement comes from scored outcomes, not from intention.

Where they part

Where positions part.

  • On scope: a general epistemic virtue, or a technique for resolvable forecasting questions.
  • On what to optimize: calibration alone, or the full decomposition including resolution.
  • On transfer: whether tournament training changes everyday judgment.

Limits

What the construct cannot do.

  • It cannot apply to unresolvable claims: no outcome, no score, and most claims about meaning or value have none.
  • It cannot be assessed on a single judgment: calibration is a property of a long run, and calling one claim well-calibrated is a category error.
  • It cannot supply resolution: a perfectly calibrated judge can be entirely uninformative.
  • It cannot detect a well-chosen question set, which is how a good score is most easily manufactured.

Criticisms

Standing objections, at strength.

  • That it quietly restricts what counts as a claim: the questions calibration can score are a narrow and unrepresentative subset of the ones that matter.
  • That the forecasting results are tournament findings whose generalization is assumed rather than demonstrated.
  • That numerical confidence can perform rigour: attaching a percentage to an intuition makes it look measured without making it measured.
  • That the construct rewards the temperamentally cautious and penalizes domains where being wrong is how progress happens.

Common misreadings

Four.

  • Calibrated as a compliment: the word used to praise a careful manner, when it names a measured property nobody has measured.
  • Single-claim calibration: assessing one statement as well-calibrated, which the concept does not permit.
  • Percentages as rigour: numeric confidence on unresolvable claims is decoration.
  • Calibration without resolution: hedging toward the base rate scores well and says nothing.

What remains outside this idea

The boundary.

  • That a judge is right: calibration concerns the honesty of confidence, not the truth of claims.
  • That anyone in this record is calibrated: no scored record of resolvable predictions exists here, and the construct's own standard forbids asserting it without one.
  • That a well-calibrated reader is a good one: informativeness is a separate property the score decomposes out.

References for further reading

Primary Source

Glenn W. Brier, 'Verification of Forecasts Expressed in Terms of Probability', Monthly Weather Review 78, no. 1 (1950): 1–3.

Secondary Source

Philip E. Tetlock and Dan Gardner, Superforecasting: The Art and Science of Prediction (Crown, 2015).