This research draft develops a Wilsonian renormalization-group (RG) view of neural-network layer convergence. It connects three ideas that have usually been discussed separately: Heavy-Tailed Self-Regularization (HTSR), the SETOL statistical-mechanics theory, and optimizer-induced spectral flow.
The central proposal is that strong trained layers approach a scale-balanced spectral boundary near α ≈ 2. Statistically, this is where learning reaches the edge of non-self-averaging. In RG power counting, it is the marginal point where equal logarithmic spectral bands contribute comparable energy. The same construction explains the SETOL trace-log condition through the retained-coordinate Jacobian and motivates RG-aware optimizer corrections such as WW-PGD.
Research draft: Charles H. Martin, PhD, Calculation Consulting, June 2026. | Read the draft PDF | GitHub staging copy | WeightWatcher GitHub
Deep learning is normally described as loss minimization. This paper argues that a second object
is equally important: the internal spectral organization of each layer. For a dense layer
with weights W, define the layer correlation operator
The eigenvectors of X identify learned input directions; its eigenvalues measure the
gain assigned to those directions. During training, many layers move from a random-matrix-like
spectrum toward a heavy-tailed spectrum. The proposed RG description organizes that movement.
Strong trained layers often develop heavy-tailed empirical spectral densities, and many good checkpoints sit near α ≈ 2.
SETOL connects layer quality to the empirical spectrum and selects a retained Effective Correlation Space (ECS) with a trace-log scale condition.
Decimate tiny-eigenvalue noise, retain correlated modes, remove arbitrary scale, and test whether the normalized spectral shape approaches a stable endpoint.
The paper argues that two diagnostics that were originally developed for different purposes actually point to the same converged layer state. HTSR fits the heavy-tailed part of the empirical spectral density and finds that well-trained layers often sit near \(\alpha\approx 2\). SETOL independently selects an effective correlation space and tests whether its normalized retained spectrum lies on the trace-log gauge slice, \(\sum_{i=1}^{\widetilde M}\log \widetilde\lambda_i \approx 0\). The proposed fixed-point picture is that both conditions hold simultaneously for a good layer.
Representative endpoint diagnostics from the paper. Left: the HTSR power-law fit gives
\(\alpha \approx 2\). Right: the trace-log-normalized ECS satisfies
\(\sum_{i=1}^{\widetilde M}\log \widetilde\lambda_i \approx 0\). The theory proposes that
strong layers satisfy both conditions at convergence.
The microscopic optimizer updates every entry of W, but the RG description does not
attempt to reproduce every weight-space detail. Instead it follows the retained empirical
spectral density. In shorthand,
One operational RG step has four parts: remove tiny-eigenvalue noise, retain the correlated ECS, rescale the retained operator, and test the remaining spectral energy and shape. The fixed-point object is not a raw matrix. It is an equivalence class of normalized retained spectra.
The RG-flow schematic below summarizes how the paper interprets optimizer training as spectral evolution. The dashed blue line marks the ECS or trace-log boundary; the dashed red line marks the start of the MLE-fitted power-law tail. During useful training, these two independently selected cutoffs move toward one another. Near convergence they align, and the fitted slope approaches \(\alpha\approx 2\).
Schematic spectral diagnostics during optimizer training and overfitting. The flow starts from a
Gaussian random-matrix-like layer, passes through increasingly heavy-tailed structured states, and
can either converge near the proposed critical set or overshoot into two different overfitting modes.
In short, the image expresses the paper's main RG claim: training drives layer spectra away from the Gaussian fixed point toward a marginal heavy-tailed boundary. Good training stops near the aligned HTSR/SETOL critical set; overtraining pushes the spectrum into trap-dominated or very-heavy-tailed regimes.
Operational spectral RG flow. Training removes weak modes, concentrates useful correlations in an
ECS, and can approach a balanced heavy-tailed shape near α ≈ 2.
This is a layerwise fixed-point picture. A multilayer network need not have one global matrix fixed point. Each matrix-like layer has its own retained spectrum and must reach a compatible endpoint with the other layers.
A macroscopic observable is self-averaging when relative sample-to-sample fluctuations vanish
as the system grows. For an observable OM, the paper uses the relative variance
When \(\mathcal{V}_M(O) \to 0\), one sufficiently large sample becomes representative of the ensemble. When it stays finite, different data samples, random seeds, or training runs can remain measurably different. In a layer spectrum, this can happen through two mechanisms:
Self-averaging allows concentration bounds to tighten with scale. Non-self-averaging leaves a finite
normalized width, so a mean-centered bound need not become informative.
The leading retained spectral observable is the first moment of the normalized ECS spectrum, interpreted as a spectral energy:
Set \(x = \log \widetilde\lambda\). Since \(d\widetilde\lambda = \widetilde\lambda\,d\log\widetilde\lambda\), the energy carried by one logarithmic spectral scale is
Energy per logarithmic band decreases toward the largest eigenvalues. The layer can be stable, but the learned tail may be too weak or underdeveloped.
Every fixed-ratio spectral band carries comparable energy. This is the proposed marginal, scale-balanced boundary.
Energy grows toward the dominant tail. A few large modes can carry an order-one fraction of the total and make the layer unstable or non-self-averaging.
The derivative argument can be written as a finite-shell calculation. Take one multiplicative spectral shell \([\Lambda,b\Lambda]\), with fixed \(b>1\), and define the band energy
For a power-law tail \(\widehat\rho_R(\widetilde\lambda)=C\widetilde\lambda^{-\alpha}\),
The second expression is independent of \(\Lambda\). Moving the same fixed-ratio shell up or down the tail does not change its energy contribution. That is the cleanest form of the marginality claim.
The paper uses the shell gain itself as a running spectral coupling,
For an exact power law, the leading spectral beta function and scaling dimension are
| Regime | Scaling dimension | Band-energy flow | Learning interpretation |
|---|---|---|---|
| α > 2 | yE < 0 |
Gain decreases with spectral scale. | Stable or self-averaging, but potentially undertrained or weakly correlated. |
| α = 2 | yE = 0 |
Gain is constant across logarithmic bands. | Marginal boundary; maximal stable structured variation before tail domination. |
| α < 2 | yE > 0 |
Gain grows toward the largest modes. | Dominant-tail sector; correlation traps and overfitting risk. |
The trace-log condition is not inserted only as a normalization trick. It arises from the volume change
of the retained coordinate map. After an SVD/ECS decimation, let WR contain the
retained right-singular directions and map retained coordinates a to layer outputs:
Because WR can be rectangular, its ordinary determinant is not the relevant
object. The induced retained-volume factor comes from the Gram determinant:
Therefore the positive retained Jacobian is
Apply the same calculation to the normalized retained operator \(\widetilde X_R\). Requiring a unit retained volume gives
The trace-log is the additive log-volume coordinate removed by the gauge choice.
This distinction is essential: the trace-log fixes scale; it is not itself the flowing spectral energy. The energy and tail shape remain physical observables after the redundant scale coordinate is removed.
Wilsonian RG acts on equivalence classes rather than every microscopic coordinate. In the retained spectral description, two important motions can be redundant.
The update changes the retained product scale and leaves the chosen gauge slice, even when the normalized spectral shape is nearly unchanged.
An orthogonal similarity transformation rotates retained coordinates while preserving the full eigenvalue multiset and the trace-log.
The trace-log correction removes one normal direction, but a similarity orbit lies tangent to the trace-log surface. Consequently, Tr log X̃R = 0 is necessary but not sufficient for a complete RG quotient.
There is also a network-level caveat. A rotation that is isospectral for one layer is not automatically redundant for the full neural network: it changes the coordinates sent to the next layer. It can be removed only when adjacent layers transform compatibly or a functional test verifies that the network output is unchanged.
The paper uses “fixed point” operationally. It does not claim that one finite checkpoint is an exact closed-form solution of the complete network dynamics. A candidate fixed point is a normalized retained spectral shape that remains stable after final decimation and trace-log testing.
The supervised optimizer remains primary. RG-aware corrections should act on a completed optimizer proposal, remain local and damped, and fall back to the original update whenever the correction damages the loss trajectory or destabilizes the retained support.
Muon approximately orthogonalizes matrix-valued updates, replacing a strongly anisotropic update spectrum by one with nearly equal active singular values. In the RG interpretation, this can suppress excessive anisotropic scale motion. However, Muon does not select an ECS, compute the layer-state trace-log normal, or determine which isospectral directions are truly redundant for the full network.
WW-PGD is a first-pass experiment that tests the proposal more directly. A base optimizer proposes a finite layer matrix. WW-PGD then selects a working spectral window between the fitted power-law start and the trace-log ECS start, reshapes the rank-ordered tail toward an α ≈ 2 target, restores the selected trace-log, and blends the result back into the optimizer trajectory.
It does not implement the complete sampled isospectral quotient. It is an approximate, falsifiable test of the RG optimizer hypothesis.
The present experiments are deliberately small. They use a three-layer MNIST MLP and track fitted WeightWatcher exponents across training. The results should be read as proof-of-concept evidence, not as a universal optimizer benchmark.
In the reported run, AdamW drives one layer below α = 2 early, while Muon follows a slower spectral route
and eventually moves the tested layer exponents toward the proposed boundary.
WW-PGD changes the layerwise spectral trajectories and pulls the tested layers toward the α ≈ 2
region.
The spectral correction preserves approximately comparable test accuracy in the reported five-run
experiment.
Starting from the selected overfit checkpoint, the preliminary WW-PGD run recovers substantially more
held-out accuracy than resumed AdamW or Muon.
The broader conjecture is that different learning algorithms may follow different microscopic update rules while approaching similar coarse-grained spectral structure. The proposed universal observables would be measured near a candidate endpoint, not read directly from the raw parameterization.
Candidate universality: different nonlinear learners can follow different microscopic trajectories yet
approach a common coarse-grained critical surface and spectral endpoint.
Neural networks provide the direct case because every matrix-like layer has an explicit correlation operator. For boosted trees and other nonlinear learners, one must first define an effective correlation operator—for example, a Gram matrix built from prediction increments—before the same spectral questions become testable.
Universality here does not mean that all algorithms compute the same function. It means that successful learners may share comparable coarse-grained spectral organization: stable bulk structure, controlled dominant modes, and a marginal balance between weak learning and overconcentration.
The paper’s value is therefore not only explanatory. It defines a concrete measurement program: track α, ECS-tail alignment, trace-log residuals, trap counts, and tail concentration across training; then ask whether generalization improves near the marginal boundary and degrades when the tail becomes dominant.
The GitHub PDF is a temporary public copy while the manuscript is being prepared for archival release. Replace the staging link with the final arXiv or journal URL when available.