ArXiv · 2026
We develop a rigorous mean-field theory for transformer networks that captures two large-scale limits inherent in the architecture: the number of tokens N→∞ in the input sequence and the number of attention heads H→∞ in each layer. It further considers an infinite number of layers leading to a time-continuous formulation. The resulting framework couples two interacting mean-field objects: a token distribution μₜ∈P(Rᵈ), which evolves through network depth t∈[0,T] according to a McKean–Vlasov transport equation, and the attention-parameter distribution ρₛ∈P(Θ), that evolves through training time s≥0 according to a Wasserstein gradient flow of the empirical risk, with optional entropic or Tikhonov regularization. We establish a comprehensive analytical foundation for the system coupling the transformer and the training dynamics; in particular, we establish global well-posedness of the resulting nonlinear Fokker–Planck system describing the coupled mean-field and training dynamics. Beyond well-posedness, we connect the mean-field formulation to optimization. For shallow, single-layer attention models, we prove exponential convergence to the entropy-regularized global optimum under a log-Sobolev condition. For genuinely deep, compositional transformers, we establish local linear convergence under a Neural Tangent Kernel non-degeneracy condition.
Try inveni