Empirical scaling laws say that loss falls as a power of parameters, data, and compute. The spectral-RG view explains why: training drives each layer toward a compatible spectral fixed surface; perturbations relax exponentially in RG time; and RG time grows logarithmically with resources. Exponentials in RG time therefore appear as power laws in resource space.
Research draft: Charles H. Martin, PhD, Calculation Consulting, July 2026 · Related: A Spectral Renormalization-Group View of Learning
Neural scaling laws are the macroscopic signature of finite-resource approach to a joint layerwise spectral fixed manifold. The fitted loss exponents are set by the slowest irrelevant spectral modes that remain visible to the loss. Exact or near marginality makes this approach unusually slow and produces logarithmic corrections, crossover, and drifting effective exponents.
Conventional scaling laws directly relate resource counts to terminal loss:
Here, P is total trainable parameter count, D is the amount of effective independent training data, and C is total training compute. These are external control parameters. They do not uniquely specify what the network learned.
P, D, C, optimizer, data mix
{ρℓ(λ), αℓ, ECSℓ, Tr log X̃ᴿ, Aℓ, Φᴷ, …}
L, error, transfer, robustness
The spectral-RG theory inserts the missing internal state. Resources determine how far the layer spectra travel toward their fixed surfaces; the resulting spectral state determines the observable loss.
For a matrix-like layer ℓ with weights \(W_\ell\), define the layer correlation operator
The eigenvalues of \(X_\ell\) describe the gains assigned to learned directions, while the empirical spectral density records how those gains are distributed. A useful layer state is
| State variable | Interpretation |
|---|---|
| ρ̂ℓ,ᴿ(λ) | The normalized retained empirical spectral density; the most complete layerwise state variable in this construction. |
| αℓ | The fitted or local heavy-tail exponent, with the candidate scale-balanced boundary near αℓ ≈ 2. |
| ℛᴱᶜˢ,ℓ | The retained Effective Correlation Space: the spectral sector carrying learned correlations after weak-noise decimation. |
| Tᴿ,ℓ = Tr log X̃ᴿ,ℓ | The retained log-volume coordinate. The ERG gauge condition is Tᴿ,ℓ ≈ 0. |
| Aℓ | Alignment between the independently selected ECS boundary and the fitted heavy-tail boundary. |
| Φᴷ,ℓ | The fraction of retained spectral energy carried by the largest K modes; a dominant-tail burden. |
| Mₜᵣ,ℓ | An effective participation count for the modes carrying trace or spectral energy. |
The full network state is the coupled collection
The distinction that matters:
P, D, and C are control coordinates. The spectra and their low-dimensional diagnostics are state coordinates. A resource law that omits the state can be predictive, but it cannot by itself explain which internal organization produced the prediction.
A multilayer network does not need one universal raw-matrix fixed point. Each layer may approach its own normalized retained spectral endpoint,
with a layer-specific retained shape, scale-fixed log volume, dominant-mode burden, and coupling to neighboring layers. The network endpoint is therefore a compatible joint manifold
Compatibility matters: an isospectral rotation that is redundant inside one layer need not be redundant for the full network unless adjacent layers transform consistently or the network function is unchanged.
Expand each layer state around its endpoint using spectral scaling operators \(\mathcal O_{\ell a}\):
Near the joint fixed manifold, the resource-scale RG flow is linear to leading order:
Stable irrelevant modes have \(\omega_{\ell a}>0\). Small \(\omega_{\ell a}\) means slow convergence. An exactly marginal linear direction has \(\omega_{\ell a}=0\) and must be resolved by nonlinear terms.
The central resource assumption is that increasing model size, effective data, or compute extends the available RG flow only logarithmically:
Substituting these relations into exponential relaxation gives
Assume terminal excess loss is a smooth readout of the joint spectral displacement. If the first nonzero coupling of mode \((\ell,a)\) occurs at order \(m_{\ell a}\ge 1\), then
Along a compute trajectory, this becomes a sum of power laws:
At sufficiently large scale, the slowest decaying mode with nonzero excitation and nonzero loss coupling dominates:
Why a single global exponent can emerge:
Each layer may have a different fixed point and many decay modes. The network still exhibits one asymptotic exponent because the slowest loss-visible mode acts as a spectral bottleneck for the whole model.
The spectral RG construction defines the first-moment shell coupling
For a power-law tail \(\widehat\rho_R(\widetilde\lambda)\sim\widetilde\lambda^{-\alpha}\),
Energy per logarithmic spectral band decreases toward the largest modes. The tail is stable but may be weak or underdeveloped.
Equal fixed-ratio spectral bands carry comparable first-moment energy. This is the candidate marginal, scale-balanced boundary.
Energy increases toward the dominant tail. A few modes or correlation traps can carry an order-one burden.
Important distinction:
The shell derivative \(d/d\log\widetilde\lambda\) compares eigenvalue scales inside one checkpoint. It is not the resource-flow derivative \(d/d\log P\), \(d/d\log D\), or \(d/d\log C\). Therefore the loss exponents p, q, and r are not equal to 2 − α. The condition α ≈ 2 identifies the candidate fixed surface; the stability spectrum around that surface determines the macroscopic scaling exponents.
The numerical exponents are therefore predicted by
Treat finite capacity and finite effective data as two cutoff fields,
Let these fields have RG dimensions \(y_P>0\) and \(y_D>0\), and let the leading loss observable have scaling dimension \(x_L>0\). RG covariance gives
Choose \(b=P^{1/y_P}\). Then
This crossover law is more fundamental than a sum of two fitted powers. It predicts that large runs in the same spectral universality class depend on the single dimensionless coordinate
When x → ∞, data are abundant relative to model capacity. If F(x) → A, then
When x → 0, the model is large relative to its independent data. If F(x) ∼ Bx⁻ᑫ, then
The minimal crossover function with both asymptotes is
Substitution recovers the familiar additive law:
In this view, the Kaplan–Chinchilla form is the leading matched approximation to an RG crossover function, not the fundamental microscopic law. Corrections naturally include additional powers, mixed terms, and logarithms:
In the unique-data regime, approximate training compute by
Minimizing
under the compute constraint gives the balance condition
The compute-optimal frontier is therefore
The optimal excess loss obeys
What equal parameter and data scaling requires:
If p = q, then Pₒₚₜ ∝ C¹ᐟ² and Dₒₚₜ ∝ C¹ᐟ², so Dₒₚₜ/Pₒₚₜ is constant. But α ≈ 2 does not imply p = q. Equal allocation exponents require an additional equality of the capacity and data RG dimensions, yᴾ = yᴰ, or an equivalent dynamical symmetry.
Starting from the general crossover law \(\Delta L=P^{-p}F(D/P^{p/q})\), fixed compute selects a constant optimal crossover coordinate \(x_\star\). Therefore
The allocation exponents are thus a consequence of RG crossover scaling itself; the additive law supplies one convenient interpolation and the corresponding amplitudes.
If a linear stability eigenvalue vanishes, the leading flow may begin quadratically:
Then
Because \(\tau\propto\log C\), exact marginality produces
This is the RG analogue of critical slowing down: the restoring rate becomes very small near marginality, and at exact marginality the linear relaxation time diverges. Calling it physical critical slowing down strictly requires measuring a training-time relaxation scale and showing that it diverges near the spectral critical manifold.
If several layerwise modes remain visible over the accessible scale range,
The measured local exponent is
It changes with scale until the slowest mode dominates. Near-degenerate modes, marginal logarithms, architecture transitions, repetition-induced instabilities, or a change in the loss-visible bottleneck all produce crossover. A straight log–log segment is therefore a finite-window effective law unless the spectral state confirms that the same fixed-point basin and the same leading mode remain active.
A useful theory must predict more than the existence of a fitted power law. The spectral-RG formulation makes several sharper tests.
For a grid of parameter and data budgets, plot
Runs in the same spectral universality class should collapse onto a single curve \(Y=F(X)\). Compute-optimal runs should lie near one constant coordinate \(X=X_\star\).
import weightwatcher as ww
watcher = ww.WeightWatcher(model=model)
details = watcher.analyze(
plot=False,
detX=True, # trace-log / ERG diagnostics
randomize=True, # correlation-trap diagnostics
)
# Track these quantities checkpoint by checkpoint and layer by layer:
# alpha, detX_val, rand_num_spikes, retained support, tail concentration, ...
Decisive benchmark:
Can small-run spectral trajectories predict the large-run loss exponents and the optimal P–D frontier more accurately than direct loss-only extrapolation? A positive result would convert the scaling exponents from fitted constants into measured properties of the spectral stability matrix.
The theory is therefore stronger than the statement “critical systems have power laws,” but it is not complete until the resource-to-RG-time map and the relevant stability eigenvalues are measured or derived. The key unknowns are \(\kappa_P,\kappa_D,\kappa_C\), the joint layerwise eigenvalues \(\omega_{\ell a}\), and the loss-coupling orders \(m_{\ell a}\).
Spectral-RG scaling proposition:
Suppose trained layers approach a compatible SETOL/HTSR spectral fixed manifold; terminal excess loss is a smooth functional of perturbations around it; and finite parameter count and finite effective data act as scaling fields with dimensions yᴾ and yᴰ. Then
The minimal crossover approximation gives \(L=L_\infty+AP^{-p}+BD^{-q}\). Under \(C=\chi PD\),
The exponents p and q are set by the leading loss-visible perturbations around the joint layerwise fixed manifold. The condition α ≈ 2 identifies marginal first-moment shell balance; it does not itself equal the loss-scaling exponent. Exact marginality adds logarithmic corrections and an RG form of critical slowing down.
In one sentence: scaling laws exist because exponential relaxation toward layerwise spectral fixed points is observed through resource coordinates whose natural RG time is logarithmic.