A shallow network with activation is a finite sum of ridge functions, each constant along a hyperplane and varying only in one direction. Letting the sum become an integral turns the network into a continuous object whose “coefficients” are a function on the space of directions. The ridgelet transform is the analysis operator of that representation. It stands to networks as the wavelet transform stands to signals, a reconstruction pair governed by an admissibility condition.

Fix an activation and its dual analysing profile . For the ridgelet transform integrates against the ridge atoms indexed by a direction and a bias :

The reconstruction rebuilds as a continuous-width network, a superposition of ridge functions weighted by :

Property. Admissibility

Reconstruction holds when the pair is admissible, i.e. the constant is finite and nonzero. This is the exact analogue of the wavelet admissibility condition. A network with activation can represent iff an analysing profile pairs with it through this one-dimensional integral. The bias plays the role of a wavelet’s translation and that of a combined rotation-and-scale.

Property. Factoring through the Radon transform

factors as a one-dimensional wavelet transform applied along each direction of the Radon transform of ; the ridge geometry decouples the radial variable from the directional variable . Consequently regularity of transverse to hyperplanes controls the decay of in , and the choice of directions lives naturally on a sphere rather than on a flat parameter space.

Remark. Categorical reading

Sonoda’s reading places the transform inside a functorial picture: the family of networks indexed by a continuous parameter is a functor, the time-evolution of the represented function under a group action is a natural transformation between two such functors, and passing to the frequency side, where the ridge atoms diagonalise, is a coordinate change realised as a natural isomorphism, with the Fourier transform as the mediating iso. The admissibility integral is then the naturality square evaluated at the identity. This is best read as a program. It is precise for the affine and rotation groups and conjectural for the nonlinear reparametrisations that deep networks induce. See functor and, for the transverse geometry of directions, riemannian-geometry.

References

  • E. J. Candes, “Ridgelets: Theory and Applications” (PhD thesis, Stanford, 1998)
  • S. Sonoda & N. Murata, “Neural Network with Unbounded Activation Functions is Universal Approximator” (Applied and Computational Harmonic Analysis, 2017)
  • S. Sonoda, on the functorial/categorical reading of the ridgelet transform