A statistical manifold is a dual-flat space carrying a Fisher metric, a pair of connections, and a canonical divergence, and it is usually built by hand from an exponential family. Is that whole package the image of a categorical construction, some functor whose output carries the geometry, with a parameter of the construction rather than of the answer?

What a construction would have to produce

The target is not a bare manifold but a bundle of coupled data: a Riemannian (Fisher) metric, a torsion-free connection and its metric dual , dual affine charts , and a canonical divergence generating all of it. A satisfying construction should emit these jointly and functorially, not glue them together after the fact.

Is there a functor whose image is the α-statistical manifold?

Find a category of “statistical models”, perhaps built over Markov categories where conditioning and sufficiency are already morphism-level notions, and a functor into a category of manifolds-with-dual-connection such that the exponential family structure maps to the dual-flat geometry and the choice of is a parameter of the functor. Then information-geometry would be, literally, the value of a construction, with the duality realised as an involution on the construction rather than a coincidence of formulas.

Two footholds

Does the divergence come from an enrichment or a Chentsov argument?

Chentsov’s theorem already characterises the Fisher metric and the -connections as the unique structures invariant under sufficient statistics, that is, under Markov morphisms. That reads as a naturality and uniqueness statement waiting for a functorial home. Separately, the canonical divergence is a Bregman divergence of the potential . Is recoverable as an enrichment, a cost or quantale structure, on the model category, so that the divergence is the hom-object and the metric its infinitesimal? If either foothold holds, “the manifold is a construction” stops being a slogan.

This is the information-geometry counterpart of asking, for convex-flow-matching, which operation selects a family. There the datum is an interpolation and the output a family, whereas here the datum is a category of models and the output the whole -geometry. A paper is planned.

References

  • Amari & Nagaoka, “Methods of Information Geometry” (2000)
  • Cencov (Chentsov), “Statistical Decision Rules and Optimal Inference” (1982)
  • Fritz, “A synthetic approach to Markov kernels, conditional independence and theorems on sufficient statistics” (2020), arXiv:1908.07021
  • Cho & Jacobs, “Disintegration and Bayesian inversion via string diagrams” (2019)
  • Fritz, Gonda, Perrone & Rischel, “Representable Markov categories and comparison of statistical experiments in categorical probability” (2023)