A Spectral Renormalization-Group View of Learning

This research draft develops a Wilsonian renormalization-group (RG) view of neural-network layer convergence. It connects three ideas that have usually been discussed separately: Heavy-Tailed Self-Regularization (HTSR), the SETOL statistical-mechanics theory, and optimizer-induced spectral flow.

The central proposal is that strong trained layers approach a scale-balanced spectral boundary near α ≈ 2. Statistically, this is where learning reaches the edge of non-self-averaging. In RG power counting, it is the marginal point where equal logarithmic spectral bands contribute comparable energy. The same construction explains the SETOL trace-log condition through the retained-coordinate Jacobian and motivates RG-aware optimizer corrections such as WW-PGD.

Main claim:
  • α > 2: stable, but the learned heavy tail can be too weak.
  • α ≈ 2: scale-balanced, maximally structured before dominant-tail takeover.
  • α < 2: tail modes or correlation traps can dominate and increase overfitting risk.
  • Tr log X̃R = 0: removes arbitrary retained-volume scale while preserving spectral shape.

Research draft: Charles H. Martin, PhD, Calculation Consulting, June 2026.  |  Read the draft PDF  |  GitHub staging copy  |  WeightWatcher GitHub


1 · The Theory in One Pass

Deep learning is normally described as loss minimization. This paper argues that a second object is equally important: the internal spectral organization of each layer. For a dense layer with weights W, define the layer correlation operator

\[ X = \frac{1}{N}W^{\mathsf T}W, \qquad \rho_{\mathrm{tail}}(\lambda) \sim \lambda^{-\alpha}. \]

The eigenvectors of X identify learned input directions; its eigenvalues measure the gain assigned to those directions. During training, many layers move from a random-matrix-like spectrum toward a heavy-tailed spectrum. The proposed RG description organizes that movement.

HTSR supplies the observation

Strong trained layers often develop heavy-tailed empirical spectral densities, and many good checkpoints sit near α ≈ 2.

SETOL supplies the effective space

SETOL connects layer quality to the empirical spectrum and selects a retained Effective Correlation Space (ECS) with a trace-log scale condition.

RG supplies the flow picture

Decimate tiny-eigenvalue noise, retain correlated modes, remove arbitrary scale, and test whether the normalized spectral shape approaches a stable endpoint.

1.1 · HTSR and SETOL identify the same layerwise endpoint

The paper argues that two diagnostics that were originally developed for different purposes actually point to the same converged layer state. HTSR fits the heavy-tailed part of the empirical spectral density and finds that well-trained layers often sit near \(\alpha\approx 2\). SETOL independently selects an effective correlation space and tests whether its normalized retained spectrum lies on the trace-log gauge slice, \(\sum_{i=1}^{\widetilde M}\log \widetilde\lambda_i \approx 0\). The proposed fixed-point picture is that both conditions hold simultaneously for a good layer.

Representative layer showing alpha approximately 2 and trace-log approximately zero Representative endpoint diagnostics from the paper. Left: the HTSR power-law fit gives \(\alpha \approx 2\). Right: the trace-log-normalized ECS satisfies \(\sum_{i=1}^{\widetilde M}\log \widetilde\lambda_i \approx 0\). The theory proposes that strong layers satisfy both conditions at convergence.
Layerwise endpoint test: a converged layer is a candidate critical layer when the independently measured HTSR and SETOL criteria line up: the retained tail is close to \(\rho(\lambda)\sim\lambda^{-2}\), and the normalized ECS lies on the trace-log gauge slice. This is the operational meaning of the paper's claim that every well-trained layer should approach \(\alpha = 2\) together with \(\operatorname{Tr}\log \widetilde X_R \approx 0\).

2 · What Actually Flows?

The microscopic optimizer updates every entry of W, but the RG description does not attempt to reproduce every weight-space detail. Instead it follows the retained empirical spectral density. In shorthand,

\[ W_{t+1} = \mathcal{O}_t[W_t], \qquad \mathcal{R}_{\mathrm{RG}}[\rho_t(\lambda)] \longrightarrow \rho^{\star}_{\mathrm{tail}}(\lambda). \]

One operational RG step has four parts: remove tiny-eigenvalue noise, retain the correlated ECS, rescale the retained operator, and test the remaining spectral energy and shape. The fixed-point object is not a raw matrix. It is an equivalence class of normalized retained spectra.

2.1 · Reading the spectral RG-flow schematic

The RG-flow schematic below summarizes how the paper interprets optimizer training as spectral evolution. The dashed blue line marks the ECS or trace-log boundary; the dashed red line marks the start of the MLE-fitted power-law tail. During useful training, these two independently selected cutoffs move toward one another. Near convergence they align, and the fitted slope approaches \(\alpha\approx 2\).

Schematic spectral diagnostics during optimizer training and overfitting Schematic spectral diagnostics during optimizer training and overfitting. The flow starts from a Gaussian random-matrix-like layer, passes through increasingly heavy-tailed structured states, and can either converge near the proposed critical set or overshoot into two different overfitting modes.

Panels 1–4: the intended training route

  1. Gaussian fixed point: no stable ECS or heavy-tailed tail is present yet.
  2. Weakly heavy-tailed: structure begins to form and the ECS cutoff appears before a mature power-law tail.
  3. Moderately heavy-tailed: correlated learning develops, and the spectrum carries a clearer heavy-tailed sector.
  4. Converged candidate critical set: the ECS boundary and fitted tail start align, while the layer sits near \(\alpha\approx 2\).

Panels 5–6: two ways to overfit

  1. Correlation-trap overfitting: a few trapped modes peel away and dominate even though the bulk tail may still look reasonable.
  2. Very heavy-tailed overfitting: the whole retained tail crosses into \(\alpha<2\), so dominant modes carry too much spectral weight and the layer becomes non-self-averaging.

In short, the image expresses the paper's main RG claim: training drives layer spectra away from the Gaussian fixed point toward a marginal heavy-tailed boundary. Good training stops near the aligned HTSR/SETOL critical set; overtraining pushes the spectrum into trap-dominated or very-heavy-tailed regimes.

Operational RG flow as decimation, rescaling, and controlled spectral energy Operational spectral RG flow. Training removes weak modes, concentrates useful correlations in an ECS, and can approach a balanced heavy-tailed shape near α ≈ 2.

This is a layerwise fixed-point picture. A multilayer network need not have one global matrix fixed point. Each matrix-like layer has its own retained spectrum and must reach a compatible endpoint with the other layers.


3 · Why α = 2 Is Special

3.1 · Statistical meaning: the self-averaging boundary

A macroscopic observable is self-averaging when relative sample-to-sample fluctuations vanish as the system grows. For an observable OM, the paper uses the relative variance

\[ \mathcal{V}_M(O) = \frac{\operatorname{Var}(O_M)}{|\mathbb{E}[O_M]|^2}. \]

When \(\mathcal{V}_M(O) \to 0\), one sufficiently large sample becomes representative of the ensemble. When it stays finite, different data samples, random seeds, or training runs can remain measurably different. In a layer spectrum, this can happen through two mechanisms:

Self-averaging and non-self-averaging distributions and relative variances Self-averaging allows concentration bounds to tighten with scale. Non-self-averaging leaves a finite normalized width, so a mean-centered bound need not become informative.

3.2 · RG meaning: equal energy per logarithmic scale

The leading retained spectral observable is the first moment of the normalized ECS spectrum, interpreted as a spectral energy:

\[ \widetilde E_R = \frac{1}{M_R}\operatorname{Tr}\widetilde X_R = \int \widetilde\lambda\,\widehat\rho_R(\widetilde\lambda) \,d\widetilde\lambda. \]

Set \(x = \log \widetilde\lambda\). Since \(d\widetilde\lambda = \widetilde\lambda\,d\log\widetilde\lambda\), the energy carried by one logarithmic spectral scale is

\[ \frac{d\widetilde E_R}{d\log\widetilde\lambda} = \widetilde\lambda^2\widehat\rho_R(\widetilde\lambda) \sim \widetilde\lambda^{2-\alpha}. \]

α > 2

Energy per logarithmic band decreases toward the largest eigenvalues. The layer can be stable, but the learned tail may be too weak or underdeveloped.

α = 2

Every fixed-ratio spectral band carries comparable energy. This is the proposed marginal, scale-balanced boundary.

α < 2

Energy grows toward the dominant tail. A few large modes can carry an order-one fraction of the total and make the layer unstable or non-self-averaging.


4 · RG Power Counting in Detail

The derivative argument can be written as a finite-shell calculation. Take one multiplicative spectral shell \([\Lambda,b\Lambda]\), with fixed \(b>1\), and define the band energy

\[ G_E(\Lambda,b) = \int_{\Lambda}^{b\Lambda} \widetilde\lambda\,\widehat\rho_R(\widetilde\lambda) \,d\widetilde\lambda. \]

For a power-law tail \(\widehat\rho_R(\widetilde\lambda)=C\widetilde\lambda^{-\alpha}\),

\[ G_E(\Lambda,b) = \frac{C\left[(b\Lambda)^{2-\alpha}-\Lambda^{2-\alpha}\right]}{2-\alpha}, \qquad \alpha\neq 2, \] \[ G_E(\Lambda,b)=C\log b, \qquad \alpha=2. \]

The second expression is independent of \(\Lambda\). Moving the same fixed-ratio shell up or down the tail does not change its energy contribution. That is the cleanest form of the marginality claim.

4.1 · Running coupling, beta function, and scaling dimension

The paper uses the shell gain itself as a running spectral coupling,

\[ g_E(\widetilde\lambda) = \frac{d\widetilde E_R}{d\log\widetilde\lambda} = \widetilde\lambda^2\widehat\rho_R(\widetilde\lambda). \]

For an exact power law, the leading spectral beta function and scaling dimension are

\[ \beta_E(g_E) = \frac{dg_E}{d\log\widetilde\lambda} = (2-\alpha)g_E, \qquad y_E=2-\alpha. \]
Regime Scaling dimension Band-energy flow Learning interpretation
α > 2 yE < 0 Gain decreases with spectral scale. Stable or self-averaging, but potentially undertrained or weakly correlated.
α = 2 yE = 0 Gain is constant across logarithmic bands. Marginal boundary; maximal stable structured variation before tail domination.
α < 2 yE > 0 Gain grows toward the largest modes. Dominant-tail sector; correlation traps and overfitting risk.
Important qualification: a power law by itself does not uniquely select α = 2, and this tree-level counting is not a complete proof of a Wilsonian fixed point. The special claim is narrower: α = 2 makes the spectral-energy coupling marginal, and the measured ECS/tail alignment provides the empirical fixed-point candidate.

5 · The Retained-Coordinate Jacobian and the Trace-Log Law

The trace-log condition is not inserted only as a normalization trick. It arises from the volume change of the retained coordinate map. After an SVD/ECS decimation, let WR contain the retained right-singular directions and map retained coordinates a to layer outputs:

\[ y(a)=\frac{1}{\sqrt N}W_Ra, \qquad \frac{\partial y}{\partial a}=\frac{1}{\sqrt N}W_R. \]

Because WR can be rectangular, its ordinary determinant is not the relevant object. The induced retained-volume factor comes from the Gram determinant:

\[ \det\!\left[ \left(\frac{W_R}{\sqrt N}\right)^{\!\mathsf T} \left(\frac{W_R}{\sqrt N}\right) \right] = \det X_R. \]

Therefore the positive retained Jacobian is

\[ J_R=\sqrt{\det X_R}=\prod_{i\in R}\sqrt{\lambda_i}, \] \[ \log J_R = \frac{1}{2}\log\det X_R = \frac{1}{2}\operatorname{Tr}\log X_R. \]

Scale-invariant ERG gauge

Apply the same calculation to the normalized retained operator \(\widetilde X_R\). Requiring a unit retained volume gives

\[ \widetilde J_R\simeq 1 \quad\Longleftrightarrow\quad \operatorname{Tr}\log\widetilde X_R\simeq 0. \]

The trace-log is the additive log-volume coordinate removed by the gauge choice.

5.1 · What is fixed, and what is still allowed to flow?

Fixed by the gauge

  • The retained product scale.
  • The covariance-volume coordinate.
  • The arbitrary multiplicative size of the retained operator.

Still free to flow

  • The tail slope and fitted α.
  • The first moment and band-energy distribution.
  • The ECS support and dominant-tail burden.
  • The normalized spectral shape.

This distinction is essential: the trace-log fixes scale; it is not itself the flowing spectral energy. The energy and tail shape remain physical observables after the redundant scale coordinate is removed.


6 · Redundant Directions and Quotient Spectral Flow

Wilsonian RG acts on equivalence classes rather than every microscopic coordinate. In the retained spectral description, two important motions can be redundant.

Trace-log or volume drift

The update changes the retained product scale and leaves the chosen gauge slice, even when the normalized spectral shape is nearly unchanged.

Isospectral coordinate motion

An orthogonal similarity transformation rotates retained coordinates while preserving the full eigenvalue multiset and the trace-log.

\[ \widetilde X_R \mapsto O^{\mathsf T}\widetilde X_RO, \qquad O(\epsilon)=e^{\epsilon A}, \qquad A^{\mathsf T}=-A, \] \[ \delta\widetilde X_R=[\widetilde X_R,A]. \]

The trace-log correction removes one normal direction, but a similarity orbit lies tangent to the trace-log surface. Consequently, Tr log X̃R = 0 is necessary but not sufficient for a complete RG quotient.

There is also a network-level caveat. A rotation that is isospectral for one layer is not automatically redundant for the full neural network: it changes the coordinates sent to the next layer. It can be removed only when adjacent layers transform compatibly or a functional test verifies that the network output is unchanged.

Physical RG motion: after redundant scale and validated coordinate motion are removed, the remaining update can redistribute spectral mass, change the fitted α, alter the ECS support, or change how much energy is carried by the dominant tail.

7 · The Operational Fixed-Point Test

The paper uses “fixed point” operationally. It does not claim that one finite checkpoint is an exact closed-form solution of the complete network dynamics. A candidate fixed point is a normalized retained spectral shape that remains stable after final decimation and trace-log testing.

One practical RG step

  1. Compute the layer spectrum.
  2. Remove tiny-eigenvalue noise.
  3. Select a retained ECS.
  4. Apply Frobenius rescaling.
  5. Measure or restore the trace-log gauge.
  6. Compare the retained shape and energy with the previous step.

Candidate endpoint signature

  • The SETOL ECS aligns with the independently fitted HTSR tail.
  • The retained normalized ESD is stable.
  • Tr log X̃R ≈ 0.
  • α ≈ 2.
  • A truncated-SVD layer restricted to the ECS preserves generalization.
\[ \mathcal{R}^{\star}_{\mathrm{ECS}} \simeq \mathcal{T}^{\star}_{\mathrm{PL}}, \qquad \rho_{\mathrm{tail}}(\widetilde\lambda)\sim\widetilde\lambda^{-2}, \qquad \operatorname{Tr}\log\widetilde X_R\simeq 0. \]

8 · What the Theory Says About Optimizers

The supervised optimizer remains primary. RG-aware corrections should act on a completed optimizer proposal, remain local and damped, and fall back to the original update whenever the correction damages the loss trajectory or destabilizes the retained support.

8.1 · Muon as a crude approximation

Muon approximately orthogonalizes matrix-valued updates, replacing a strongly anisotropic update spectrum by one with nearly equal active singular values. In the RG interpretation, this can suppress excessive anisotropic scale motion. However, Muon does not select an ECS, compute the layer-state trace-log normal, or determine which isospectral directions are truly redundant for the full network.

8.2 · WW-PGD as a finite spectral retraction

WW-PGD is a first-pass experiment that tests the proposal more directly. A base optimizer proposes a finite layer matrix. WW-PGD then selects a working spectral window between the fitted power-law start and the trace-log ECS start, reshapes the rank-ordered tail toward an α ≈ 2 target, restores the selected trace-log, and blends the result back into the optimizer trajectory.

What WW-PGD controls

  • One retained log-volume coordinate.
  • The rank-order shape of a selected tail window.
  • A soft homotopy toward the α ≈ 2 target.
  • A post-correction residual used to weaken or skip unsafe corrections.

It does not implement the complete sampled isospectral quotient. It is an approximate, falsifiable test of the RG optimizer hypothesis.


9 · Preliminary Experimental Evidence

The present experiments are deliberately small. They use a three-layer MNIST MLP and track fitted WeightWatcher exponents across training. The results should be read as proof-of-concept evidence, not as a universal optimizer benchmark.

9.1 · AdamW versus Muon

MNIST AdamW versus Muon accuracy and layerwise alpha trajectories In the reported run, AdamW drives one layer below α = 2 early, while Muon follows a slower spectral route and eventually moves the tested layer exponents toward the proposed boundary.

9.2 · WW-PGD on ordinary training

Layerwise alpha values for AdamW baseline and WW-PGD WW-PGD changes the layerwise spectral trajectories and pulls the tested layers toward the α ≈ 2 region.
MNIST test accuracy for AdamW baseline and WW-PGD The spectral correction preserves approximately comparable test accuracy in the reported five-run experiment.

9.3 · Repairing an anti-grokking checkpoint

WW-PGD recovery of test accuracy from an anti-grokking checkpoint Starting from the selected overfit checkpoint, the preliminary WW-PGD run recovers substantially more held-out accuracy than resumed AdamW or Muon.
What these experiments support: optimizer geometry changes the spectral route, and a targeted finite spectral correction can improve the tested overfit trajectories while keeping the supervised optimizer in control. They do not yet prove convergence to a universal RG fixed point.

10 · Universality Beyond One Architecture

The broader conjecture is that different learning algorithms may follow different microscopic update rules while approaching similar coarse-grained spectral structure. The proposed universal observables would be measured near a candidate endpoint, not read directly from the raw parameterization.

RG trajectories for different learning algorithms approaching a candidate fixed point Candidate universality: different nonlinear learners can follow different microscopic trajectories yet approach a common coarse-grained critical surface and spectral endpoint.

Neural networks provide the direct case because every matrix-like layer has an explicit correlation operator. For boosted trees and other nonlinear learners, one must first define an effective correlation operator—for example, a Gram matrix built from prediction increments—before the same spectral questions become testable.

Universality here does not mean that all algorithms compute the same function. It means that successful learners may share comparable coarse-grained spectral organization: stable bulk structure, controlled dominant modes, and a marginal balance between weak learning and overconcentration.


11 · What Is Established, and What Remains a Proposal?

Empirical or operational ingredients

  • WeightWatcher measures layerwise ESDs and fitted heavy-tail exponents.
  • Many strong models exhibit heavy-tailed layer spectra.
  • Correlation traps can be tested by randomized-spectrum outliers.
  • The ECS, trace-log residual, tail burden, and α can be tracked across checkpoints.
  • The reported Muon and WW-PGD experiments can be reproduced and falsified.

Theoretical claims still under test

  • That α ≈ 2 is a universal layerwise RG endpoint.
  • That the self-averaging and marginality boundaries coincide broadly.
  • That common optimizers converge to the proposed spectral fixed point.
  • That a complete quotient optimizer can remove redundant directions safely.
  • That the same universality class extends to boosted trees and other learners.

The paper’s value is therefore not only explanatory. It defines a concrete measurement program: track α, ECS-tail alignment, trace-log residuals, trap counts, and tail concentration across training; then ask whether generalization improves near the marginal boundary and degrades when the tail becomes dominant.