Dynamic Depth-wise Magnitude & Angular Superposition
TL;DRHidden-state norms, variance grow by orders of magnitude with depth, a “curse of depth” treated as a pathology. We find it encodes an emergent depth code in the norm weights γ, and make it explicit. LayerRoPE is a RoPE-style depth encoding in the residual stream. It improves compute scaling, depth scaling up to 512 layers, and widens the learning-rate basin. Surprisingly, it does this by learning to amplify residual variance (and regulating Block I/O), rather than dampening it.
A Transformer exposes exactly one learned gain per layer directly on the residual path: the RMSNorm weight γ. That makes it the most direct place for any depth-conditional code the network may wish to learn.
Across 16 open-weight LLMs from 9 families (dense, mixture-of-experts, parallel-branch and hybrid linear-attention; Pre-, Peri- and Post-Norm), we find that the MLP-block γ carries exactly such a depth-conditioned code:
Magnitude is depth-conditioned. The mean and upper percentiles of γ shift systematically with depth, rising or falling with clear depth conditioning in most models.
Direction is depth-conditioned. We observe that the normalized gain \gamma_\ell / \lVert \gamma_\ell \rVert also drifts steadily with depth. MDS projections, shown below, order the layers along a coherent structured arc, showing a clear directional drift with the layers. This is not a rigid rotation, however. See1
Where RoPE encodes a token’s position by rotating pairs of query and key features, LayerRoPE encodes a block’s layer index by rotating and scaling a shared gain. Instead of storing L independent vectors γℓ, every layer’s gain is generated from one shared γ ∈ ℝᵈ per site and a single depth-conditioned complex factor:
LayerRoPE can thus be simply interpreted as a depth-conditioned scaling and rotation of γ in the complex plane. Reading the real gain as d/2 complex pairs \tilde\gamma_j = \gamma_{2j} + i\,\gamma_{2j+1}, the magnitude scales all pairs equally through a log-linear schedule (the shared trajectory found above), while the rotation spreads a per-layer base angle across pairs with RoPE’s fixed spectrum, b₀ being the base frequency.
LayerRoPE sits at four sites (the read and the write of every attention and MLP block), so it regulates how each block reads from and writes to the residual stream, rather than the stream itself. Four scalars per site replace O(Ld) per-layer gain parameters with O(d).
import torch, torch.nn as nn class LayerRoPE(nn.Module): """Depth-conditioned gains for one site (e.g. the attention read). One shared γ and four scalars replace L per-layer gain vectors.""" def __init__(self, d, n_layers, base=100.0, beta_init=-0.5): super().__init__() self.gamma = nn.Parameter(torch.ones(d)) # shared γ self.a_mag = nn.Parameter(torch.zeros(())) self.b_mag = nn.Parameter(torch.tensor(beta_init)) self.a_rot = nn.Parameter(torch.zeros(())) self.b_rot = nn.Parameter(torch.tensor(beta_init)) j = torch.arange(d // 2) self.register_buffer("freq", base ** (-2 * j / d)) # RoPE spectrum self.register_buffer("log_l", torch.log(torch.arange(1.0, n_layers + 1))) def forward(self): # input-independent: once per pass r = self.a_mag + self.b_mag * self.log_l theta = torch.exp(self.a_rot + self.b_rot * self.log_l) ang = theta[:, None] * self.freq c = torch.polar(torch.exp(r)[:, None].expand_as(ang), ang) g = torch.view_as_complex(self.gamma.view(-1, 2)) return torch.view_as_real(g * c).flatten(1) # One LayerRoPE module per site sites = [LayerRoPE(d, n_layers) for _ in range(4)] attn_read, attn_write, mlp_read, mlp_write = (site() for site in sites) # normalize(h) = h / RMS(h), i.e. RMSNorm without its own learned weight: for layer, block in enumerate(blocks): h = h + attn_write[layer] * block.attn(normalize(h) * attn_read[layer]) h = h + mlp_write[layer] * block.mlp(normalize(h) * mlp_read[layer])
We train LLaMA-style models from 58M to 1.3B parameters on C4 at 4× the Chinchilla budget, up to 107B tokens. Every method, from Pre-, Post- and Peri-Norm to Layer-Norm Scaling (LNS) and LayerRoPE, gets its own learning rate, swept until the optimum is bracketed and extrapolated to 1.3B; LayerRoPE’s own hyperparameters stay fixed across scales.
LayerRoPE attains the lowest loss at every scale with the steepest compute exponent.
To isolate the curse of depth, we fix a narrow backbone (width 128) and vary only the number of layers, from 48 to 512 (18M–111M parameters), adding DeepNet, designed for very deep Transformers, as a baseline. All baselines diverge with depth (Post-Norm from 192 layers, Peri-Norm and LNS degrade beyond 96, and Pre-Norm stops improving after 256). DeepNet improves with depth, but remains well above LayerRoPE throughout.
LayerRoPE reaches the lowest loss, with a margin that grows with depth; an exhaustive per-depth learning-rate sweep also confirms our finding.
Sweeping the peak learning rate over two to three orders of magnitude, with and without LayerRoPE on both backbones, changes the loss curve in three ways:
With no tuning of their recipes, LayerRoPE carries over to two architectures far from the language-model ladder: a looped latent language model and Vision Transformers.
Parcae applies a recurrent block of layers T times to a latent state between an input and an output block. We train it at 140M, 370M and 770M parameters with its official recipe, beside a non-looped GPT of the same size, and apply LayerRoPE at every layer (each layer’s depth index is its position in the network, independent of the loop iteration), with magnitude slopes initialised at zero. Parcae + LayerRoPE is the best model on nearly every metric: lower validation perplexity, markedly lower Lambada perplexity at 140M and 370M, and higher average downstream accuracy.
140M11.2B tokens | 370M29.6B tokens | 770M61.6B tokens | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Val ppl ↓ | LMB ppl ↓ | Avg ↑ | Val ppl ↓ | LMB ppl ↓ | Avg ↑ | Val ppl ↓ | LMB ppl ↓ | Avg ↑ |
| GPT | 20.28 | 105.6 | 29.9 | 15.01 | 37.2 | 40.6 | 12.80 | 20.8 | 47.0 |
| Parcae | 18.66 | 81.2 | 33.0 | 14.22 | 32.5 | 43.4 | 12.49 | 18.6 | 50.1 |
| + LayerRoPE | 18.36 | 67.3 | 34.9 | 14.05 | 28.0 | 44.0 | 12.09 | 18.4 | 50.6 |
We train ViT-T and ViT-S on ImageNet-1k for 300 epochs with the DeiT recipe and apply LayerRoPE naively at every block, conditioning the gains γ of the pre-norm LayerNorms while keeping their centring and bias, with the magnitude slope initialised at zero. LayerRoPE improves both ViTs on every metric.
Prior remedies such as Layer-Norm Scaling and Peri-Norm damp the residual stream directly. LayerRoPE’s learned schedule does the opposite.
LayerRoPE does not shrink the residual pipe; it widens it, while regulating the blocks that read from and write to it.
How? LayerRoPE lets the residual stream grow in variance and norm, and instead regulates it at the inputs and outputs of each computational block.
Measuring just the learned magnitude gains, we see:
LayerRoPE learns to attenuate what each block reads from the residual stream.
LayerRoPE learns to amplify each block's write backs into the residual stream.
LayerRoPE’s behaviour, of residual-stream amplification with regulation of each block’s I/O, is therefore a learned property.
Learning stability & depth scaling, our results suggest, depend less on how large the residual stream grows than on whether each block can scale what it reads and writes to its depth.
We find that LayerRoPE’s depth-conditioned I/O enables this regulation of the computational block, while amplifying the variance of the residual pipe.
This work was partly supported by NSF EFRI Award 2317706 and NSF CAREER Awards 2047556/2326491. The views and conclusions contained herein are those of the authors and should not be interpreted as representing any sponsor’s official policies or endorsements. We gratefully acknowledge use of the research computing resources of the Empire AI Consortium, Inc, with support from Empire State Development of the State of New York, the Simons Foundation, and the Secunda Family Foundation. This work was also supported in part by compute credits provided by Modal Labs through the Modal for Academics program.