ArXiv · 2026
Deep neural networks exhibit regular macroscopic behavior despite highly nonlinear dynamics in vast parameter spaces. We develop a statistical-mechanical description of learning directly in function space, treating parameter configurations as microscopic realizations and functions with their dynamical operators as macroscopic variables. For mean-squared loss, the exact error dynamics are governed by the learning operator M=JJ^∗. Combining the dynamical Boltzmann weight of the conditional stochastic dynamics with the parameter-space density of states, whose local curvature defines a statistical operator B, and integrating over local fluctuations yields Φ_fluc(M;B)=σ_ξ²/2logdet(M⁻¹+B)+const. At fixed spectrum, this term is rotationally stationary when [M,B]=0, is minimized by pairing large eigenvalues of M with small eigenvalues of B, and generates a local restoring contribution against rotational mismatch. For ReLU-type function spaces under mild stable statistical conditions, B=σ_ξ²L^∗K L, where L measures coarse-grained second-order structure. Thus the low-B sector corresponds, up to bounded anisotropy of K, to low structural curvature, implying a preference for faster relaxation along smooth, data-adaptive directions. These results identify function space as a natural macroscopic level for studying stable collective organization in learning.
Try inveni