Every Transformer Lives on a Sphere
I first met normalization through Goodfellow, Bengio, and Courville's wonderfully clear treatment of it: standardize the activations a layer receives, then give the network a learned scale and shift so it can recover any useful range.[13] The LayerNorm paper makes the next move feel almost trivial. Take BatchNorm's recipe and turn it ninety degrees: instead of comparing one feature across examples, compare all the features inside one example.[1]
That explanation is excellent. It is enough to implement LayerNorm and to understand why a sequence model would rather not make one sentence depend on whichever other sentences happened to share its batch. But I kept wanting to ask whether there was more hiding in the equation. Was LayerNorm only tidying up unruly numbers, or was it changing the kind of object a hidden state could be?
Then I drew it.
The statistical recipe turned into a geometric operation: delete one direction, forget one magnitude, and leave the token on the intersection of a plane and a sphere. The learned affine parameters may warp that surface afterwards, and a real implementation adds a small \(\epsilon\), but the central picture survives. A token entering LayerNorm with thousands of coordinates leaves with only a direction in a smaller space.
That picture connects a surprising decade of work. RMSNorm asks whether the deleted direction was worth deleting. A 2023 result argues that the geometry changes what attention can express. nGPT puts most of the Transformer on hyperspheres deliberately. Work from 2026 finds an exact dead direction created by LayerNorm and begins treating residual updates as movement on a manifold. Meanwhile, normalization-free Transformers ask whether the sphere was essential in the first place.
The title is intentionally provocative.Not literally every Transformer. Ordinary LayerNorm and RMSNorm have direct spherical interpretations before their learned affine maps. nGPT makes that constraint architectural. Dynamic Tanh and other normalization-free systems are the useful counterexamples that keep the title honest. A Transformer using LayerNorm does not keep every hidden state permanently on one sphere, and a model using Dynamic Tanh need not visit a sphere at all. But spheres are now a remarkably good route through the history, mathematics, and current design of normalization.
First, which numbers share a ruler?
Normalization is easiest to understand before it becomes geometry. Imagine a small table of activations \(X\in\mathbb R^{B\times d}\). Each row is one example; each column is one learned feature. To standardize a group of numbers means to move its centre to zero and set its typical spread to one. The learned parameters \(\gamma\) and \(\beta\) can then stretch and shift the result. BatchNorm and LayerNorm mostly disagree about one question: which numbers belong to the group?
That axis choice has practical consequences. During training, BatchNorm's output for one example can change when the other examples in its minibatch change, and inference usually relies on accumulated population estimates. LayerNorm needs neither: it computes fresh statistics inside each token, so its rule is the same with a batch of one or a batch of one thousand.A convolutional footnote. In common CNN implementations, BatchNorm pools a channel over the batch and spatial locations, not merely down one tidy column. The small table isolates the essential distinction: BatchNorm shares statistics across examples; LayerNorm does not.
So far this is a story about bookkeeping axes. Now rotate the point of view. One row is not merely a list of feature values; it is a vector in a space. “Across the row” then becomes “along particular directions,” and standardization becomes geometry.
How a statistical trick became a geometry
The chronology matters because nobody began by announcing a manifold. The original 2016 LayerNorm paper moved normalization statistics from the batch to the features of one example.[1] RMSNorm removed the mean subtraction in 2019.[2] Only later did a sequence of papers make the implied geometry explicit, connect it to attention, and then turn it into an architectural principle.
Start with the boring equation
Take one token in a residual stream. Its hidden representation is a vector
\[ x= \begin{bmatrix} x_1 & x_2 & \cdots & x_d \end{bmatrix}^{\!\top} \in\mathbb R^d. \]The residual stream is the model's running workspace. Each Transformer block reads the current vector, computes an attention or MLP update, and adds that update back. The same token representation therefore accumulates many layers of changes. Its coordinates are not human-labelled features; they are a learned basis in which the network stores whatever distinctions help prediction. Normalization sits at a consequential junction: it decides what version of that workspace the next sublayer is allowed to read.
The coordinates are features, not tokens.Which axis is normalized? For the usual language-model case, LayerNorm treats the \(d\) features of each token separately from the other tokens and examples. BatchNorm instead estimates statistics using a batch dimension. This per-token independence is one reason LayerNorm fits variable-length sequence models so naturally. Ordinary LayerNorm computes the mean and variance across those \(d\) coordinates:
\[ \mu=\frac1d\sum_{i=1}^{d}x_i, \qquad \sigma^2=\frac1d\sum_{i=1}^{d}(x_i-\mu)^2. \]Temporarily ignore numerical stabilisation and the learned scale and shift. The standardized output is
\[ \operatorname{LN}_0(x)=\frac{x-\mu\mathbf1}{\sigma}, \qquad \mathbf1= \begin{bmatrix}1&1&\cdots&1\end{bmatrix}^{\!\top}. \]This is how I had always read the operation: coordinate-wise statistics. But the vector \(\mathbf1\) is a clue. Every coordinate of it moves together, which means it represents a very specific direction through the hidden space. Once we ask what subtracting \(\mu\mathbf1\) does along that direction, the equation stops looking statistical.
Subtracting the mean deletes a direction
Normalize the all-ones vector:
\[ u=\frac{\mathbf1}{\sqrt d}, \qquad u^\top u=1. \]The component of \(x\) pointing along \(u\) is its orthogonal projection
\[ \operatorname{proj}_u(x)=(u^\top x)u. \]Now expand the dot product. Because \(u\) contains \(d\) copies of \(1/\sqrt d\),
\[ u^\top x =\frac1{\sqrt d}\sum_i x_i =\sqrt d\,\mu. \]Multiplying by \(u\) again gives
\[ (u^\top x)u =\sqrt d\,\mu\frac{\mathbf1}{\sqrt d} =\mu\mathbf1. \]So the unremarkable phrase subtract the mean is exactly the operation
\[ x-\mu\mathbf1=x-\operatorname{proj}_u(x). \]LayerNorm removes the part of the token that points along the uniform direction. It is not softly discouraged or learned away. Before \(\epsilon\) and the affine map enter, that component is annihilated by construction.
The same operation can be written as a matrix:
\[ P=I-\frac1d\mathbf1\mathbf1^\top, \qquad Px=x-\mu\mathbf1. \]The identities \(P^\top=P\) and \(P^2=P\) say that \(P\) is an orthogonal projector.Why does \(P^2=P\) matter? Project once and the vector is already in the plane; projecting again changes nothing. Symmetry, \(P^\top=P\), says the projection is perpendicular rather than oblique. Together these are the algebraic fingerprints of an orthogonal projection. Every centered vector obeys
\[ \mathbf1^\top Px=0, \]so it lives in the hyperplane \(H=\mathbf1^\perp\). The original space had \(d\) independent directions. The constraint that all coordinates sum to zero leaves \(d-1\).
There is an operational way to feel that lost direction. Add the same number to every coordinate:
\[ x\longrightarrow x+c\mathbf1. \]LayerNorm cannot see the move, because
\[ P(x+c\mathbf1)=Px+cP\mathbf1=Px. \]A common offset can be arbitrarily large and the standardized result will not notice. LayerNorm has declared that translation along \(\mathbf1\) is irrelevant.
Dividing by standard deviation creates the sphere
The denominator now becomes geometric too. The variance is the average squared length of the centered coordinates:
\[ \sigma^2=\frac1d\|Px\|_2^2, \qquad \sigma=\frac{\|Px\|_2}{\sqrt d}. \]Substitute that into LayerNorm:
That single equation is the article. First project away the uniform component. Then divide by the remaining length, which keeps only orientation. Finally multiply by \(\sqrt d\), which chooses the radius. Every standardized output \(y\) therefore satisfies
\[ \mathbf1^\top y=0, \qquad \|y\|_2^2=d. \]One equation restricts \(y\) to a hyperplane. The other restricts it to a hypersphere. LayerNorm outputs lie where those two surfaces meet.[3][7]
There is also a quick statistical check on the radius. “Unit variance” across coordinates means their average squared value is one after centering. Multiply that average by the \(d\) coordinates and the sum of squares is \(d\). The Euclidean length must therefore be \(\sqrt d\). The sphere is not an extra geometric claim laid on top of the statistics; it is the statistics read as a length constraint.
This is where dimension language is easy to muddle. The centered hyperplane has dimension \(d-1\). Fixing the radius removes one more continuous degree of freedom. The standardized outputs themselves form a sphere \(S^{d-2}\) of radius \(\sqrt d\), embedded inside that hyperplane.Sphere notation. \(S^n\) names a surface of intrinsic dimension \(n\): \(S^1\) is a circle and \(S^2\) is the ordinary globe surface. The sphere \(S^{d-2}\) lives inside a \((d-1)\)-dimensional centered space, which itself sits inside \(\mathbb R^d\).
\(S^0\): two points
\(S^1\): a circle
\(S^2\): a sphere
\(S^{4094}\)
The two-dimensional shock
The first rung is worth doing by hand because it makes the information loss impossible to ignore. Let \(x=(a,b)\). Its mean is \((a+b)/2\), so centering gives
\[ x-\mu\mathbf1 =\frac{a-b}{2}(1,-1). \]No matter where \(x\) began, its centered version points along the same diagonal line. Its standard deviation is \(|a-b|/2\), so dividing leaves only the sign:
\[ \operatorname{LN}_0(a,b) = \begin{cases} (1,-1), & a>b,\\ (-1,1), & a<b. \end{cases} \]Every nonconstant point in the entire plane becomes one of two answers. The degenerate case \(a=b\) has zero variance, which is exactly why implementations need \(\epsilon\). Real Transformers use far more than two dimensions, so this is not a warning that all useful information vanishes. It is a clean demonstration of what the operator keeps: not coordinate level, not length, only which centered direction survived.
Normalization laboratory
Move the three coordinates, switch the operator, and drag the drawing to rotate it. The black arrow is the input direction, the grey arrow is the centered component, and the blue arrow is the normalized output.
mean = 0.667 output norm = √3 output sum = 0
LayerNorm does not merely normalize. It forgets.
There is a more complete way to describe what disappeared. Let
\[ x'=ax+b\mathbf1, \qquad a>0. \]Projecting removes the common offset and keeps the positive scaling:
\[ Px'=aPx. \]Normalizing then removes that scaling:
\[ \sqrt d\frac{aPx}{\|aPx\|} =\sqrt d\frac{Px}{\|Px\|}. \]Therefore
\[ \boxed{ \operatorname{LN}_0(ax+b\mathbf1) =\operatorname{LN}_0(x)} \qquad(a>0). \]An entire two-parameter family of vectors becomes one point. LayerNorm ignores common offset and overall positive scale; it retains only the direction of the centered vector. The word normalization sounds like one vector being cleaned up. The operation is more accurately collapsing an equivalence class.
Abstract algebra gives that collapse a name. Define \(x\sim y\) when \(y=ax+b\mathbf1\) for \(a>0\). LayerNorm maps the quotient—original vectors modulo those two nuisance transformations—to the sphere \(S^{d-2}\). We did not add abstraction for decoration. The quotient tells us exactly what information no later calculation can reconstruct from the standardized vector alone.
This is the mirror image of the idea in my essay Your LLM Does Group Theory Every Time It Reads a Sentence. RoPE takes a symmetry—translation by an integer number of token positions—and represents it as a rotation, preserving the action so attention can see relative position. LayerNorm takes two transformations—uniform translation and positive rescaling—and quotients them away, declaring that attention and the MLP should not see them. RoPE remembers a group action; LayerNorm deliberately forgets one. Both become much less mysterious once the Transformer's arithmetic is read as geometry.
Negative scaling is the small twist. If \(a<0\), the output flips:
\[ \operatorname{LN}_0(-x)=-\operatorname{LN}_0(x). \]Opposite directions remain antipodal points rather than becoming identical. That is why the surviving space is a sphere, not real projective space.Projective space, very briefly. Real projective space identifies a line through the origin, so \(v\) and \(-v\) represent the same point. LayerNorm preserves that sign. It forgets positive radius but still distinguishes the two orientations of the line.
Real implementations use \(\epsilon\):
\[ y=\frac{Px}{\sqrt{\|Px\|^2/d+\epsilon}}, \qquad \|y\|=\sqrt d\sqrt{\frac{\operatorname{Var}(x)}{\operatorname{Var}(x)+\epsilon}}. \]So the radius is slightly below \(\sqrt d\), approaching it whenever the variance dominates \(\epsilon\). PyTorch currently defaults to \(10^{-5}\).[11] Then the learned map \(z=\gamma\odot y+\beta\) can stretch different coordinates unequally and translate the result. The standardized vector lies on the sphere; learned per-coordinate scaling generally warps that sphere into an affine, hyperellipsoidal surface.[4]
The derivative keeps only tangent motion
The forward map says which points survive. Its Jacobian says which tiny movements survive. Let
\[ z=Px, \qquad u_z=\frac{z}{\|z\|}, \qquad y=\sqrt d\,u_z. \]Ignoring \(\epsilon\), differentiating gives
\[ J_{\mathrm{LN}} =\frac{\sqrt d}{\|z\|} \left(I-u_zu_z^\top\right)P. \]This looks denser than the forward equation but contains the same story twice. The projector \(P\) kills motion along \(\mathbf1\). The projector \(I-u_zu_z^\top\) kills motion in the current radial direction. What remains is motion tangent to the sphere.
This is the differential-geometric version of the quotient. LayerNorm removes two local directions, leaving \(d-2\) directions along which the token can move. The prefactor \(\sqrt d/\|z\|\) rescales those tangent movements, so the operation also changes gradient magnitudes; geometry does not replace the usual optimization story, it sharpens it.
RMSNorm keeps the direction LayerNorm deletes
RMSNorm begins with the root mean square
\[ \operatorname{RMS}(x) =\sqrt{\frac1d\sum_i x_i^2} =\frac{\|x\|}{\sqrt d}. \]Again ignoring \(\epsilon\) and the learned scale,
\[ \operatorname{RMSNorm}_0(x) =\sqrt d\frac{x}{\|x\|}. \]Put the equations beside each other and the architectural difference becomes one letter:
LayerNorm
\[ \operatorname{LN}_0(x) =\sqrt d\frac{Px}{\|Px\|} \]\(\mathbb R^d\to\mathbf1^\perp\to S^{d-2}\)
forgets offset and positive scale
RMSNorm
\[ \operatorname{RMSNorm}_0(x) =\sqrt d\frac{x}{\|x\|} \]\(\mathbb R^d\to S^{d-1}\)
forgets positive scale only
RMSNorm does not slice the sphere with the zero-mean plane. It lands on the full \(S^{d-1}\), keeping the uniform component if one exists. Its 2019 paper described the same distinction in statistical language: retain rescaling invariance, hypothesize that re-centering invariance is dispensable, and remove the mean computation.[2] The interactive lab above lets you see the difference. In RMSNorm mode, the black and blue arrows always point the same way. In LayerNorm mode, projection usually rotates the direction before fixing its radius.
Neither operator is simply “more spherical.” Both forget radius. LayerNorm additionally chooses one privileged axis and removes it. That asymmetry is exactly what later work found interesting.
Why might attention want this geometry?
Brody, Alon, and Yahav argued in 2023 that LayerNorm contributes to attention's expressivity, not only its numerical stability.[3] Their two geometric mechanisms are unusually clean.
First, standardized keys obey \(\mathbf1^\top k_i=0\). A query parallel to \(\mathbf1\) therefore gives every key the same score:
\[ q=c\mathbf1 \quad\Longrightarrow\quad q^\top k_i=c\mathbf1^\top k_i=0. \]Equal logits become equal softmax weights. Projection gives attention a cheap route to uniform averaging—a useful primitive for tasks such as counting a majority.
Second, suppose one key is a strict convex combination of others, \(k_n=\sum_{i<n}\lambda_i k_i\), with nonnegative \(\lambda_i\) summing to one. For every query,
\[ q^\top k_n=\sum_{i<n}\lambda_i q^\top k_i. \]Its score is an average of other scores, so it cannot be strictly largest. The key is unselectable. Distinct vectors with the same norm cannot be strict interior convex combinations of one another: such a combination falls inside the sphere. Equal-radius scaling therefore keeps every normalized key on the boundary and potentially able to win.
The 90-degree mystery
The 2026 EACL study by Gupta and colleagues asked whether models actually carry much information along \(\mathbf1\). They passed one million Wikipedia tokens through seven decoder-only LLMs, capturing representations around every normalization layer. Five used LayerNorm—GPT-2 XL, GPT-Neo 1.3B, Pythia 1.4B, GPT-J 6B, and Pythia 6.9B—and two used RMSNorm—Llama-2 7B and Llama-3 8B.[7]
The diagnostic was just the angle \(\theta=\angle(h,\mathbf1)\). For four of the five LayerNorm models, hidden states were already close to \(90^\circ\) before mean subtraction; GPT-2 XL was the exception. More surprisingly, both Llama models stayed close to orthogonal even though RMSNorm never explicitly removes that component. Controlled models from 70M to 1B parameters, trained with either operator on OpenWebText, developed the same broad pattern during training.
That last qualification matters.High-dimensional surprise. Two random isotropic vectors in large dimension are usually close to \(90^\circ\). The interesting observation is not that ninety degrees can occur, but that it persists across layers, checkpoints, operators, and controlled training—and that RMSNorm models do not drift into the direction LayerNorm would have removed. The experiment is evidence that explicit centering is often redundant in the tested regimes, not proof that means never matter in every architecture or distribution.
The missing direction leaves an algebraic fingerprint
Projection does not merely make a nice picture. Before the affine map, every output satisfies \(\mathbf1^\top y=0\). Its centered covariance \(C=\operatorname{Cov}(y)\) must therefore obey
\[ C\mathbf1=0. \]The uniform vector is an exact kernel direction: variance along it is zero. Restore the learned scaling \(D_\gamma=\operatorname{diag}(\gamma)\) and shift \(\beta\):
\[ z=D_\gamma y+\beta, \qquad \operatorname{Cov}(z)=D_\gamma C D_\gamma. \]The constant shift disappears from covariance. Now choose \(v=D_\gamma^{-1}\mathbf1\). Then
\[ \operatorname{Cov}(z)v =D_\gamma C\mathbf1=0. \]So the affine map moves the dead direction from \(\mathbf1\) to the normalized inverse-scale vector
\[ \boxed{v^\star=\frac{\gamma^{-1}}{\|\gamma^{-1}\|}}. \]Shirodkar and Narayanan tested this 2026 prediction on 14 pretrained Transformers from 160M to 35B parameters. The predicted vector aligned with the measured bottom singular direction on all nine LayerNorm models; five RMSNorm models lacked the corresponding universal kernel.[8] “Subtract the mean” had become a parameter-readable singular direction in billion-parameter networks.
Not a neuron, and not necessarily a useless model. A covariance kernel is a direction with no centered activation variation at that location. The striking part is that architecture alone predicts it. Training may create additional low-variance directions, but this one is already written into LayerNorm's projector.
Then researchers made the metaphor literal
Once hidden states have fixed norm, Euclidean addition becomes suspicious. The update \(h_{\ell+1}=h_\ell+\Delta h\) generally leaves the sphere. A manifold-aware update instead begins with a tangent vector \(v_\ell\in T_{h_\ell}S^{d-1}\) and returns through a retraction:
\[ h_{\ell+1}=R_{h_\ell}(v_\ell). \]nGPT, published at ICLR 2025, normalizes embeddings, hidden states, and vectors inside attention and MLP blocks onto unit hyperspheres. The authors describe each layer as moving the token's representation across the surface and reported reaching equivalent accuracy in 4–20 times fewer optimization steps in their experiments.[5] A 2026 training recipe then scaled the idea to hybrid Mamba-2/Transformer mixture-of-experts models up to 14B total parameters; its 14B comparison reached the same validation loss with roughly half as many training tokens.[9]
Perhaps the sphere is not the essential part
A good geometric story can become too satisfying. Dynamic Tanh is the necessary interruption. The CVPR 2025 paper Transformers without Normalization replaced sample-wise statistics with the pointwise operation
\[ \operatorname{DyT}(x)=\gamma\odot\tanh(\alpha x)+\beta. \]No mean, no variance, no spherical projection. Yet the authors reported matching or exceeding normalized counterparts across language, vision, diffusion, and speech settings. Their starting observation was that trained normalization layers often produce tanh-like input-output curves.[6] Later normalization-free work such as Derf continued the argument with other bounded pointwise functions.[12]
That leaves two plausible readings. Hyperspherical constraints may provide a particularly useful package of invariances and optimization geometry. Or the larger benefit may be controlled activation scale and well-shaped gradients, with the sphere only one way to obtain them. Current evidence supports both as productive design directions, not one final verdict.
The direction was the representation
I began with an equation that seemed to standardize coordinates. It ended by changing the ontology of a hidden state. An input vector was not preserved as a shorter or cleaner version of itself. Its uniform component was removed, its positive radius was forgotten, and the equivalence class that remained was represented by a point on a sphere.
That single change of viewpoint explains the dimension count, the tangent-space Jacobian, the contrast with RMSNorm, a route to uniform attention, the convex-hull result, and an exact covariance kernel. It also explains why hyperspherical Transformers do not feel like an arbitrary new metaphor. They make explicit a geometry that normalization had been quietly imposing for years.
But Dynamic Tanh keeps the ending open. Maybe the sphere is the mechanism. Maybe it is an especially elegant scaffold. Either way, “subtract the mean and divide by the standard deviation” is no longer a sufficient mental model.
Normalization chooses what can be forgotten. Geometry tells us what is left.
References and links
- J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer Normalization,” 2016.
- B. Zhang and R. Sennrich, “Root Mean Square Layer Normalization,” NeurIPS 2019.
- S. Brody, U. Alon, and E. Yahav, “On the Expressivity Role of LayerNorm in Transformers' Attention,” Findings of ACL 2023.
- P. M. Riechers, “Geometry and Dynamics of LayerNorm,” 2024.
- I. Loshchilov, C.-P. Hsieh, S. Sun, and B. Ginsburg, “nGPT: Normalized Transformer with Representation Learning on the Hypersphere,” ICLR 2025.
- J. Zhu, X. Chen, K. He, Y. LeCun, and Z. Liu, “Transformers without Normalization,” CVPR 2025.
- A. Gupta, A. Ozdemir, C. Gong, and G. Anumanchipalli, “Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm,” Findings of EACL 2026.
- T. P. Shirodkar and P. J. Narayanan, “Algebraic Dead Directions in LayerNorm Transformers: A Forward-Pass-Only Diagnostic at LLM Scale,” 2026.
- I. Loshchilov and B. Ginsburg, “Training nGPT,” 2026.
- J. Zhang, C.-F. Su, Y.-J. Huang, and M.-T. Sun, “Sphere Retraction Normalizations,” 2026.
- PyTorch LayerNorm documentation, including the \(\epsilon\) and element-wise affine conventions.
- M. Chen, T. Lu, J. Zhu, M. Sun, and Z. Liu, “Stronger Normalization-Free Transformers,” 2025.
- I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning, Chapter 8: “Optimization for Training Deep Models,” MIT Press, 2016.
- S. Ioffe and C. Szegedy, “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” ICML 2015.