~/projects

A Transport-Based Geometry of Belief-Cost

What it costs a digital twin to change its mind about a finite world.

Laurent Caraffa

Université Gustave Eiffel, LASTIG, IGN-ENSG · Preprint · 2026 · arXiv:2606.21585

Two postulates, one hyperbolic geometry. A belief is a probability distribution over the world, here a Gaussian: a point with mean and width . Certainty () is the boundary and sits at infinite cost-distance, so it stays out of reach; the space has negative (hyperbolic) curvature. Changing a belief reshapes the Gaussian, at a cost measured as a length in nats (the base-e unit of information). The strip below the two panels shows the belief along each route from to : the geodesic (violet) widens the Gaussian then re-narrows it; the straight chord (blue) keeps the width fixed and costs more. Drag a belief in either model.

Abstract

A belief is a probability density, and changing it has a cost. This project turns that cost into a geometry: the cost of moving from one belief to another is a distance, built from optimal transport (the least effort to reshape one density into another) reweighted by Fisher information (the sharpness of the belief). The framework holds for any probability density, and rests on two modest assumptions: revision cost is a scalar price on transport, and that price is uniform, with one nat costing the same length everywhere.

For any belief:

On the Gaussian (location-scale) family, indexed by a mean and a width:

A finite agent and its belief

A finite machine observes a fixed world through noisy sensors, so its belief is a probability density over states (the Bayes posterior), always keeping a positive width.

Certainty (a point mass) is ruled out by two facts, both read from the Fisher information : observation (Cramér–Rao: the width shrinks only as ) and physics (Landauer: erasing one nat costs , so exact certainty would cost infinite energy).

The problem. Characterize the geometry of the beliefs a finite agent can hold: its distances (the cost of changing one's mind), its boundary (certainty), and its curvature.

Certainty is ruled out twice. The belief (violet) sharpens as observations accrue but stays a density; the barred Dirac spike is the excluded point mass. Below, the two limits, both reading . Drag the sliders.

The two postulates

Finiteness fixes the object (a density, certainty excluded) but leaves the geometry open. Everything below rests on two postulates, stated on the location-scale family (the Gaussian beliefs). Each is a modelling choice; the rest is forced or standard.

Postulate 0 · the transport metric

Cost is a scalar price on transport

The cost of changing a belief is a scalar price on the optimal transport of probability mass: the transport metric multiplied by a positive scalar field (a cost-height over belief space). It prices a revision; Bayes performs it. Multiplying a metric by a scalar is a conformal reweighting ( is a constant offset, here 0):

Postulate 1 · uniform price

One nat costs the same length everywhere

The price is uniform (eikonal): the metric slope of knowledge is constant across belief space. Under transport, Fisher information sets the slope of entropy rather than a distance, and uniform pricing then selects one family:

The two are ordered: P0 sets the metric and its scalar price; P1 fixes that price. Thermodynamics motivates both, but the proofs use only geometry, in nats.

The geometry the cost imposes

From the two postulates the geometry is fixed. On the location-scale leaf (the plane of the headline figure) it is hyperbolic: the cheapest revision is a curved geodesic, not the straight chord, and equal-cost steps crowd toward the boundary. That boundary is certainty (), at infinite cost-distance ().

The boundary

Certainty is unreachable

The cost to reach certainty diverges as : .

The Fisher family

Uniform pricing ⇔ Fisher

Uniform pricing (each nat the same length everywhere) is equivalent to , the Fisher family. This characterizes Postulate 1.

Stam rigidity

The Gaussian is extremal

The Stam bound makes the Gaussian the most negatively curved location-scale belief: .

All three are invariant under a change of cost unit; the value is one example.

The distance, in four formulas

Four equations take the geometry from definition to computed number. Throughout, the constants are c = 2 and e = 0.

The two ingredients: Fisher information J (the sharpness of a density) and differential entropy H (its spread in nats); knowledge is −H. For a Gaussian, .

The cost of a revision. A path of beliefs γ has a cost equal to its transport speed times the local price , integrated; the distance from A to B is the smallest such cost, in nats.

The cost metric on the Gaussian leaf. Here the transport metric is flat: the distance between two Gaussians is the Euclidean distance in , so . With , the local price becomes : the Poincaré half-plane with every length doubled. A step costs 2/σ nats per metre, which doubles each time σ halves (the cost-per-step shading behind the figures).

The distance between two Gaussian beliefs and : the standard half-plane distance along the geodesic. The worked example below uses this formula for each revision (in one go, and one sensor at a time); the chord cost uses the length element of eq. 3 along the straight segment.

Why the distance is measured in nats. On the Fisher family (, ) the metric does not depend on the unit chosen for the state. Rescale and shift the height axis, with : the transport distance scales by (optimal transport carries the state's length), while the Fisher information scales by , so the local price scales by . In the length the factor from transport and the from Fisher cancel, so every distance is unchanged. Concretely, eq. 4 depends only on the dimensionless ratio , invariant under . The units of optimal transport and of Fisher information cancel, leaving a metric with no dimension: distances are pure information, in nats. The cancellation holds precisely at the exponent in , which is also the exponent uniform pricing selects.

A worked example: the digital twin of a field

One concrete inference, end to end: a fixed field, a fixed set of noisy sensors, a kriging belief, and the cost of each revision. Every displayed number is computed in the browser, from the kriging formulas below and eqs. 1–4 above.

The problem

A field C carries a vegetation height h(x). n point sensors at positions return unbiased measurements with known Gaussian noise . The likelihood of a candidate field is ; it constrains h at the n sensor sites. A belief in the sense above is a probability distribution over the whole field, so the setup is completed by a spatial prior (the standard geostatistics choice): a Gaussian process with squared-exponential kernel (amplitude a, correlation range ℓ). With Gaussian noise the posterior is Gaussian, called kriging. Writing K for the kernel matrix at the sensors, , and for the kernel vector between x and the sensors,

The belief is this Gaussian posterior. Read at one location x₀, it is a scalar Gaussian , one point of the half-plane, where eqs. 3 and 4 apply. Two standard kriging facts matter below: given the kernel, the variance σ²(x) depends only on the sensor positions and noise levels (the measured values enter the mean alone), and it decreases weakly as sensors are added.

The figure, step by step

The digital twin of a field. Left: the field and its kriging belief. Right: the belief space of the readout at x₀, the (μ, σ) half-plane, with height in metres on the horizontal axis. Two tabs select the regime, Learn (n) and Revise (two campaigns); each panel carries its own title, gridlines and legend.

Left panel: the field

The blue curve is the kriging mean m(x); the blue band is m(x) ± 2σ(x). The dotted vertical line is the readout location x₀ (drag it, or use the x₀ slider); the orange bar there spans μ ± 2σ. In the Learn tab, thin violet curves are three sample fields drawn from the posterior. Filled orange dots are the sensors read; hollow dots are the unread pool (40 in all; the n slider selects a subset). The s and ℓ sliders set the sensor noise and the prior range; “Another dataset” redraws the noise. The toggle shows h★, the true field.

Right panel: belief space

Each point is a scalar belief about the height at x₀. The bold line σ = 0 is certainty, at infinite cost-distance. The violet shading is the cost map: each band down doubles the price per step (×2, ×4, ×8), and every band has the same cost thickness, 2 ln 2 nats. The local price of moving the mean is per metre. Markers show one nat of cost each along both routes (one per two nats on routes past 20).

Learn (n) tab

The posteriors after k = 0, 1, …, 40 sensors (posterior 0 is the prior) are joined by the orange polyline, solid up to the current n and faint beyond. The violet arc is the geodesic from posterior 0 (or the pinned A) to posterior n. The readouts give σ(x₀), the last step's cost, the step-by-step total with its ratio, the direct revision, and the whole-field information Σₖ J(xₖ) (the total sharpness summed over field locations, which rises as the belief tightens). Rings mark n = 5, 10, 20, 40: each doubling of the evidence costs a step of similar length (in the i.i.d. limit, exactly ½ ln 2 nat). “Shuffle the order” permutes the sensors; posterior 0 and posterior 40 hold still, since the order is a gauge. “Pin posterior A” sets the arc's start to the current posterior, and the readout then shows d(A→B) = d(B→A).

Revise (two campaigns) tab

Two campaigns survey the same field; campaign 2 has a declared calibration offset of +0.45 m. Kriging gives both the same width σ (the variance reads the sensor geometry alone), with readings 0.45 m apart and means ≈ 0.44 m apart at x₀: two confident beliefs that disagree. On the right panel the endpoints are A (campaign 2, violet) and B (campaign 1, blue). The solid violet curve is the cheapest revision on the Gaussian leaf: it widens (σ rises), traverses, then re-narrows. The dashed blue segment is the W₂ chord, the straight route in (μ, σ), which for Gaussians is the optimal-transport (McCann) interpolation: the Gaussian slides sideways at fixed width. At the defaults the geodesic costs nats and the chord nats (Δμ = 0.445 m, σ = 0.0886 m), a saving of 34 %. The entropy floor is 0.00 here (equal widths), yet the revision still costs ≈ 6.6 nats: cost is a length, and entropy change is one part of it.

The “budget spent b” slider (Revise tab)

b is a number of nats, spent equally on both routes; scrub it in either direction. For b > 0, a traveler sits on each route at spend b, and the “remaining” readout shows what each still owes (0 on arrival): at b = 6.61 the geodesic traveler has arrived while the chord traveler owes 3.44 nats. The same slider draws the two intermediate beliefs at x₀ as sideways Gaussians (violet widening then re-narrowing, blue staying narrow) and, on wide screens, highlights the current snapshot in the strip below: one belief per nat, 8 snapshots for the geodesic against 12 for the chord. The cheaper route spends its nats widening, and needs fewer.

Interpretation

Certainty is a boundary at infinite cost-distance. At the defaults, σ(x₀) falls from 0.65 m (the prior) to ≈ 0.09 m at n = 40. The floor diverges as σ → 0, so reaching σ = 0 costs infinite length. Every reachable belief keeps a positive width.

Every nat has one price. In the i.i.d. slice at x₀ (), doubling the evidence halves the variance and lowers the entropy by exactly ½ ln 2 nat, at a near-constant length. Kriging sits below this ideal (correlated sensors are partly redundant): each doubling n = 5 → 10 → 20 → 40 lowers the entropy by ≈ 0.2–0.3 nat. In cost coordinates the rate becomes : a fixed price per unit distance.

Adding data one sensor at a time costs more than one update. Bayes reaches the same posterior through any grouping of the data, but the costs differ: the one-sensor updates total two to three times the direct revision at n = 40, and about one and a half at the default n = 8 (the ratio grows with n). This is the triangle inequality (moving the mean costs 2/σ per metre: ≈ 3 nats at σ = 0.65 m, ≈ 22 at σ = 0.09 m). The Revise tab isolates the complementary fact: a revision with ΔH = 0 still costs ≈ 6.6 nats, and the cheapest way to pay it widens the belief first.

Scope

The right panel is the scalar slice at x₀. Every linear readout of the field (a height at a point, a mean over a plot) has a scalar Gaussian posterior on the Gaussian leaf, where each displayed quantity is a closed form from the paper. For the full field posterior , the boundary at infinite cost-distance and the entropy floor hold in every dimension. The curvature and the Stam statement are results on the two-dimensional location-scale leaf; the curvature of the full field posterior is open (Conjecture C2).

BibTeX

@misc{caraffa2026costgeometry,
  title         = {A Transport-Based Geometry of Belief-Cost:
                   an axiomatic characterization},
  author        = {Caraffa, Laurent},
  year          = {2026},
  eprint        = {2606.21585},
  archivePrefix = {arXiv},
  primaryClass  = {math.ST},
  url           = {https://arxiv.org/abs/2606.21585}
}