Classical asymptotics assume the map from parameter to distribution is a local diffeomorphism, so that the Fisher metric is nondefinite and the likelihood is locally quadratic. This fails for neural networks, mixtures, and most hierarchical models, where distinct parameters realise the same distribution and the Fisher metric degenerates on a subvariety. Watanabe’s singular learning theory is the replacement, built by resolving those singularities.

Let be a statistical model with parameter in a compact , a prior , and a true law . The model is regular when is one-to-one and the Fisher information is everywhere nondegenerate; it is singular otherwise. The object of study is the Kullback–Leibler function , whose zero set is, in the singular case, not a point but an analytic variety with self-intersections and non-normal crossings.

Definition of the real log-canonical threshold

The real log-canonical threshold (RLCT) is the largest constant such that the zeta function is holomorphic on ; equivalently is the largest pole of , with multiplicity its order. By Hironaka resolution of singularities there is a proper analytic making a normal crossing monomial, and then read off the exponents. In the regular case and .

Theorem on free-energy and generalisation asymptotics

With samples the Bayes free energy (stochastic complexity) satisfies , and the Bayes generalisation error behaves like in leading order. Thus replaces the of the Bayesian information criterion, and with strict inequality exactly at singular points. Singular models therefore generalise better than their dimension would suggest.

The RLCT is a birational invariant of the pair and is generally not the half-dimension of ; it mixes codimension with the multiplicity of the crossing. Computing it is algebraic geometry rather than asymptotic statistics. Toric resolutions and Newton polyhedra give for concrete families such as reduced-rank regression, mixtures, and small networks, whereas only bounds are known for generic architectures.

Remark on the spectral link

The Hessian of near is the empirical Fisher matrix, whose spectrum concentrates on a bulk with a large near-null part reflecting the degeneracy. This connects the RLCT to the limiting spectral distribution studied by random-matrix theory and free probability. The density of small eigenvalues controls the effective parameter count, and the free-multiplicative structure of products of layer Jacobians is a natural language for how that null part accumulates with depth. The precise dictionary between the pole of and the edge of the spectral measure is at present conjectural.

References

  • S. Watanabe, “Algebraic Geometry and Statistical Learning Theory” (Cambridge, 2009)