Two postulates, one hyperbolic geometry. A belief is a probability distribution over the
world, here a Gaussian: a point with mean
and width . Certainty
() is the boundary and
sits at infinite cost-distance, so it stays out of reach; the space has negative (hyperbolic)
curvature. Changing a belief reshapes the Gaussian, at a cost measured as a length in nats (the
base-e unit of information). The
strip below the two panels shows the belief along each route from to
: the geodesic (violet) widens the Gaussian then re-narrows it; the
straight chord (blue) keeps the width fixed and costs more.
Drag a belief in either model.
Abstract
A belief is a probability density, and changing it has a cost. This project turns that cost into a
geometry: the cost of moving from one belief to another is a distance, built from optimal transport
(the least effort to reshape one density into another) reweighted by Fisher information (the
sharpness of the belief). The framework holds for any probability density, and rests on two modest
assumptions: revision cost is a scalar price on transport, and that price is uniform, with one nat
costing the same length everywhere.
For any belief:
Certainty lies at infinite cost-distance. Every belief the agent reaches keeps a
positive width.
Uniform pricing selects Fisher information. Requiring the same price per nat everywhere
is equivalent to reweighting transport by the Fisher information
().
Distances are dimensionless, measured in nats. The units of optimal transport and of
Fisher information cancel, so distances do not depend on the units used for the world.
On the Gaussian (location-scale) family, indexed by a mean and a width:
The geometry is hyperbolic, and the Gaussian is its extreme case. The curvature is
negative and constant, and the Stam bound () makes the Gaussian
the most curved belief (). The curvature of a general belief
is left open.
A finite agent and its belief
A finite machine observes a fixed world through noisy sensors, so its belief is a probability
density over states (the Bayes posterior), always keeping a positive width.
Certainty (a point mass) is ruled out by two facts, both read from the Fisher information
: observation (Cramér–Rao: the width shrinks only as
) and physics (Landauer: erasing one nat
costs , so exact certainty would cost infinite energy).
The problem. Characterize the geometry of the beliefs a finite agent can hold:
its distances (the cost of changing one's mind), its boundary (certainty), and its curvature.
Certainty is ruled out twice. The belief (violet) sharpens as observations
accrue but stays a density; the barred Dirac spike is the excluded point mass. Below, the two
limits, both reading . Drag the sliders.
The two postulates
Finiteness fixes the object (a density, certainty excluded) but leaves the geometry open.
Everything below rests on two postulates, stated on the location-scale family (the Gaussian
beliefs). Each is a modelling choice; the rest is forced or
standard.
Postulate 0 · the transport metric
Cost is a scalar price on transport
The cost of changing a belief is a scalar price on the optimal transport of probability mass:
the transport metric multiplied by a positive scalar
field (a cost-height over belief space). It prices a revision; Bayes
performs it. Multiplying a metric by a scalar is a conformal reweighting
( is a constant offset, here 0):
Postulate 1 · uniform price
One nat costs the same length everywhere
The price is uniform (eikonal): the metric slope of knowledge is
constant across belief space. Under transport, Fisher information sets the slope of entropy
rather than a distance, and uniform pricing then selects one family:
The two are ordered: P0 sets the metric and its scalar price;
P1 fixes that price. Thermodynamics motivates both, but the proofs use only
geometry, in nats.
The geometry the cost imposes
From the two postulates the geometry is fixed. On the location-scale leaf (the
plane of the headline figure) it is
hyperbolic: the cheapest revision is a curved geodesic, not the straight chord,
and equal-cost steps crowd toward the boundary. That boundary is certainty
(), at infinite cost-distance
().
The boundary
Certainty is unreachable
The cost to reach certainty diverges as :
.
The Fisher family
Uniform pricing ⇔ Fisher
Uniform pricing (each nat the same length everywhere) is equivalent to
, the Fisher family. This characterizes Postulate 1.
Stam rigidity
The Gaussian is extremal
The Stam bound makes the Gaussian the most negatively curved location-scale belief:
.
All three are invariant under a change of cost unit; the value
is one example.
The distance, in four formulas
Four equations take the geometry from definition to computed number. Throughout, the constants
are c = 2 and e = 0.
The two ingredients: Fisher information J (the sharpness of a density) and differential entropy H
(its spread in nats); knowledge is −H. For a Gaussian, .
The cost of a revision. A path of beliefs γ has a cost equal to its transport speed times the
local price , integrated; the distance from A to B is the
smallest such cost, in nats.
The cost metric on the Gaussian leaf. Here the transport metric is flat: the
distance between two Gaussians is the Euclidean distance in
, so .
With , the local price
becomes : the Poincaré half-plane with every length doubled. A
step costs 2/σ nats per metre, which doubles each time σ halves (the cost-per-step shading behind
the figures).
The distance between two Gaussian beliefs and
: the standard half-plane distance along the geodesic.
The worked example below uses this formula for each revision (in one go, and one sensor at a time);
the chord cost uses the length element of eq. 3 along the straight segment.
Why the distance is measured in nats. On the Fisher family
(, ) the metric does not depend on the
unit chosen for the state. Rescale and shift the height axis,
with : the transport
distance scales by (optimal transport carries the state's length),
while the Fisher information scales by , so the local price
scales by . In the length
the factor
from transport and the from Fisher
cancel, so every distance is unchanged. Concretely, eq. 4 depends only on the dimensionless
ratio
,
invariant under . The units of optimal transport and
of Fisher information cancel, leaving a metric with no dimension: distances are pure
information, in nats. The cancellation holds precisely at the exponent
in , which is also the
exponent uniform pricing selects.
A worked example: the digital twin of a field
One concrete inference, end to end: a fixed field, a fixed set of noisy sensors, a kriging belief,
and the cost of each revision. Every displayed number is computed in the browser, from the kriging
formulas below and eqs. 1–4 above.
The problem
A field C carries a vegetation height h(x). n point sensors at positions
return unbiased measurements
with known Gaussian noise
. The likelihood of a candidate
field is ; it constrains h at
the n sensor sites. A belief in the sense above is a probability distribution over the whole field,
so the setup is completed by a spatial prior (the standard geostatistics choice): a Gaussian
process with squared-exponential kernel
(amplitude a, correlation range ℓ).
With Gaussian noise the posterior is Gaussian, called kriging. Writing K for the kernel matrix at
the sensors, , and
for the kernel vector between x and the sensors,
The belief is this Gaussian posterior. Read at one location x₀, it is a scalar Gaussian
, one point of the
half-plane, where eqs. 3 and 4 apply. Two standard kriging
facts matter below: given the kernel, the variance σ²(x) depends only on the sensor positions and
noise levels (the measured values enter the mean alone), and it decreases weakly as sensors are
added.
The figure, step by step
The digital twin of a field. Left: the field and its kriging belief.
Right: the belief space of the readout at x₀, the (μ, σ) half-plane, with height in metres on the
horizontal axis. Two tabs select the regime, Learn (n) and
Revise (two campaigns); each panel carries its own title, gridlines and legend.
Left panel: the field
The blue curve is the kriging mean m(x); the blue band is m(x) ± 2σ(x). The dotted vertical
line is the readout location x₀ (drag it, or use the x₀ slider); the orange bar there spans
μ ± 2σ. In the Learn tab, thin violet curves are three sample fields drawn from the posterior.
Filled orange dots are the sensors read; hollow dots are the unread pool (40 in all; the
n slider selects a subset). The s and ℓ sliders set the sensor noise and the prior range;
“Another dataset” redraws the noise. The toggle shows h★, the true field.
Right panel: belief space
Each point is a scalar belief about the
height at x₀. The bold line σ = 0 is certainty, at infinite cost-distance. The violet shading is
the cost map: each band down doubles the price per step (×2, ×4, ×8), and every band has the same
cost thickness, 2 ln 2 nats. The local price of moving the mean is
per metre. Markers show one nat of cost each along
both routes (one per two nats on routes past 20).
Learn (n) tab
The posteriors after k = 0, 1, …, 40 sensors (posterior 0 is the prior) are joined by the
orange polyline, solid up to the current n and faint beyond. The violet arc is the geodesic from
posterior 0 (or the pinned A) to posterior n. The readouts give σ(x₀), the last step's cost, the
step-by-step total with its ratio, the direct revision, and the whole-field information
Σₖ J(xₖ) (the total sharpness summed over field locations, which rises as the belief tightens).
Rings mark n = 5, 10, 20, 40: each doubling of the evidence costs a step of similar length (in
the i.i.d. limit, exactly ½ ln 2 nat). “Shuffle the order” permutes the sensors; posterior 0 and
posterior 40 hold still, since the order is a gauge. “Pin posterior A” sets the arc's start to
the current posterior, and the readout then shows d(A→B) = d(B→A).
Revise (two campaigns) tab
Two campaigns survey the same field; campaign 2 has a declared calibration offset of +0.45 m.
Kriging gives both the same width σ (the variance reads the sensor geometry alone), with readings
0.45 m apart and means ≈ 0.44 m apart at x₀: two confident beliefs that disagree. On the right
panel the endpoints are A (campaign 2, violet) and B (campaign 1, blue). The solid violet curve
is the cheapest revision on the Gaussian leaf: it widens (σ rises), traverses, then re-narrows.
The dashed blue segment is the W₂ chord, the straight route in (μ, σ), which for Gaussians is the
optimal-transport (McCann) interpolation: the Gaussian slides sideways at fixed width. At the
defaults the geodesic costs
nats and the chord nats
(Δμ = 0.445 m, σ = 0.0886 m), a saving of 34 %. The entropy floor
is 0.00 here (equal widths), yet the revision
still costs ≈ 6.6 nats: cost is a length, and entropy change is one part of it.
The “budget spent b” slider (Revise tab)
b is a number of nats, spent equally on both routes; scrub it in either direction. For
b > 0, a traveler sits on each route at spend b, and the “remaining” readout shows what each
still owes (0 on arrival): at b = 6.61 the geodesic traveler has arrived while the chord traveler
owes 3.44 nats. The same slider draws the two intermediate beliefs at x₀ as sideways Gaussians
(violet widening then re-narrowing, blue staying narrow) and, on wide screens, highlights the
current snapshot in the strip below: one belief per nat, 8 snapshots for the geodesic against 12
for the chord. The cheaper route spends its nats widening, and needs fewer.
Interpretation
Certainty is a boundary at infinite cost-distance. At the defaults, σ(x₀) falls from
0.65 m (the prior) to ≈ 0.09 m at n = 40. The floor
diverges as σ → 0, so reaching σ = 0 costs
infinite length. Every reachable belief keeps a positive width.
Every nat has one price. In the i.i.d. slice at x₀
(), doubling the evidence halves the variance and
lowers the entropy by exactly ½ ln 2 nat, at a near-constant length. Kriging sits below this ideal
(correlated sensors are partly redundant): each doubling n = 5 → 10 → 20 → 40 lowers the entropy by
≈ 0.2–0.3 nat. In cost coordinates the rate becomes
: a fixed price per unit
distance.
Adding data one sensor at a time costs more than one update. Bayes reaches the same
posterior through any grouping of the data, but the costs differ: the one-sensor updates total two
to three times the direct revision at n = 40, and about one and a half at the default n = 8 (the
ratio grows with n). This is the triangle inequality (moving the mean costs 2/σ per metre: ≈ 3 nats
at σ = 0.65 m, ≈ 22 at σ = 0.09 m). The Revise tab isolates the complementary fact: a revision with ΔH = 0 still
costs ≈ 6.6 nats, and the cheapest way to pay it widens the belief first.
Scope
The right panel is the scalar slice at x₀. Every linear readout of the field (a height at a point,
a mean over a plot) has a scalar Gaussian posterior on the Gaussian leaf, where each displayed
quantity is a closed form from the paper. For the full field posterior
, the boundary at infinite cost-distance and the
entropy floor hold in every dimension. The curvature and the
Stam statement are results on the two-dimensional location-scale leaf; the curvature of the full
field posterior is open (Conjecture C2).
BibTeX
@misc{caraffa2026costgeometry,
title = {A Transport-Based Geometry of Belief-Cost:
an axiomatic characterization},
author = {Caraffa, Laurent},
year = {2026},
eprint = {2606.21585},
archivePrefix = {arXiv},
primaryClass = {math.ST},
url = {https://arxiv.org/abs/2606.21585}
}