ArXiv · 2026
Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on why this transition occurs, the quantitative structure of when it occurs in hyperparameter space remains uncharacterized. We map the memorization-to-generalization boundary across 384 configurations of two-hidden-layer MLPs on modular arithmetic, fitting a power-law scaling relation for generalization onset time: T_grok ∝ H^(-0.27) D^(-2.04) η^(-0.50) λ^(-0.64) (R² = 0.732; 0.821 with interactions). The exponent hierarchy reveals that data complexity (D^(-2.04)) is the dominant driver of regime transition, not model capacity (H^(-0.27)): doubling data accelerates generalization by ∼4×, while doubling width yields only ∼1.2×. A sharp phase boundary at weight decay λ ≳ 1.0 separates grokking from non-grokking configurations, and weight norm trajectories show monotonic compression during the transition, consistent with implicit regularization selecting low-complexity solutions. These results provide a quantitative foundation for predicting and controlling regime transitions in overparameterized networks.
Try inveni