ArXiv · 2026
Deep learning builds complex global functions through the repeated composition of local functions. But the resulting distribution over global functions is unknown. We derive this distribution exactly for finite width and depth when the local functions are chosen at random. Surprisingly, depth first creates a sharply structured distribution, then erases this structure as probability drains into two absorbing states. This competition produces a crossover at a depth exponential in the number of inputs, beyond which additional layers mainly drive collapse. The transient structure may help explain why highly overparameterized networks can still generalize.
Try inveni