Back to blog

Why a massless vector field transforms with a gauge shift

March 21, 2026

Introduction

Gauge invariance is not a quantum idea. It is already built into classical electrodynamics.

In the classical theory the physical electric and magnetic fields are

E=−∇ϕ−∂A∂t,B=∇×A.\mathbf{E}=-\nabla \phi-\frac{\partial \mathbf{A}}{\partial t}, \qquad \mathbf{B}=\nabla\times \mathbf{A}.

The scalar potential ϕ\phi and the vector potential A\mathbf{A} are not unique. If χ(t,x)\chi(t,\mathbf{x}) is any sufficiently smooth function, then

A→A+∇χ,ϕ→ϕ−∂χ∂t\mathbf{A}\to \mathbf{A}+\nabla\chi, \qquad \phi\to \phi-\frac{\partial \chi}{\partial t}

leaves both E\mathbf{E} and B\mathbf{B} unchanged. This is the ordinary gauge invariance of Maxwell theory. In covariant notation, with

Fμν=∂μAν−∂νAμ,F_{\mu\nu}=\partial_\mu A_\nu-\partial_\nu A_\mu,

the same statement is simply

Aμ→Aμ+∂μχ,Fμν→Fμν.A_\mu \to A_\mu+\partial_\mu \chi, \qquad F_{\mu\nu}\to F_{\mu\nu}.

The invariance follows because

∂μ∂νχ−∂ν∂μχ=0.\partial_\mu\partial_\nu\chi-\partial_\nu\partial_\mu\chi=0.

So even before quantization, the potential AμA_\mu contains redundant information. Classical electrodynamics only assigns direct physical meaning to gauge-invariant quantities such as FμνF_{\mu\nu}, or equivalently E\mathbf{E} and B\mathbf{B}.

The more interesting question is why this same redundancy becomes almost unavoidable in quantum field theory. In classical electrodynamics gauge invariance may first appear as a curious non-uniqueness of the potentials. In relativistic quantum theory it becomes the mechanism that lets a local Lorentz-covariant vector field describe a massless spin-1 particle while still matching the unitary representations of the Poincare group.

There is a famous statement in quantum field theory:

if a massless spin-1 particle is described by a Lorentz-covariant field Aμ(x)A_\mu(x), then under a Lorentz transformation the field cannot transform as a pure four-vector. It must transform as a four-vector plus a gauge transformation.

In formulas, one does not have simply

Aμ(x)→ΛμνAν(Λ−1x),A_\mu(x)\to {\Lambda_\mu}^{\nu} A_\nu(\Lambda^{-1}x),

but rather

Aμ(x)→ΛμνAν(Λ−1x)+∂μΩΛ(x).A_\mu(x)\to {\Lambda_\mu}^{\nu} A_\nu(\Lambda^{-1}x)+\partial_\mu \Omega_\Lambda(x).

Why is that unavoidable?

The answer is encoded in the way one-particle states transform under spacetime symmetries. In this post I will follow the logic used by Weinberg and work through the mathematics explicitly.

Before starting, let me make two important clarifications.

1. This is about massless spin-1 particles

For a massive vector boson the little group is different and the argument below does not go through in the same way.

2. One helicity or two?

Strictly speaking, Lorentz invariance by itself does not force a massless particle to come with both helicities +1+1 and −1-1.

For the proper orthochronous Poincare group, an irreducible unitary massless representation may have a single helicity hh. If the theory is also invariant under parity, then parity flips helicity and therefore one must include both +h+h and −h-h.

So:

  • Lorentz invariance + unitarity allows one-helicity massless representations.
  • Parity invariance forces the pair (+h,−h)(+h,-h).

Parity is not in the connected Lorentz group. It belongs to the full Lorentz group O(1,3)O(1,3) as a disconnected transformation.

None of that changes the main conclusion of this post: whenever we try to describe a massless helicity-1 particle with a local Lorentz four-vector AμA_\mu, the field must transform up to a gauge shift.

Why the little group appears

In quantum mechanics the states live in a Hilbert space, and a symmetry is a transformation that preserves transition probabilities. Wigner’s theorem then says that any such symmetry must be represented on the Hilbert space by either a unitary or an antiunitary operator. Weinberg phrases the starting point in exactly this way: symmetries in quantum mechanics are implemented by operators that preserve inner products, with the allowed possibilities being unitary or antiunitary.

For continuous spacetime symmetries connected to the identity, such as ordinary Lorentz transformations and translations, one uses the unitary branch. Antiunitary symmetries appear for transformations like time reversal, but they are not the relevant case in the connected Poincare transformations considered below.

So in relativistic quantum theory one studies unitary representations of the Poincare group: Lorentz transformations plus spacetime translations.

Translations let us label one-particle states by four-momentum. We write such states schematically as

∣p,σ⟩,|p,\sigma\rangle,

where pμp^\mu is the four-momentum and σ\sigma denotes any remaining internal label, such as spin or helicity.

A Lorentz transformation changes the momentum:

pμ→Λμνpν.p^\mu \to {\Lambda^\mu}_\nu p^\nu.

So the first job of a Lorentz transformation is kinematical: it moves us from the state with momentum pp to a state with momentum Λp\Lambda p. But this does not yet tell us how the spin or helicity labels transform.

The key observation is that all momenta with the same invariant mass lie on the same Lorentz orbit. Therefore we can choose one convenient reference momentum kμk^\mu, called the standard momentum, and obtain any other momentum pμp^\mu on that orbit by some Lorentz transformation L(p)L(p):

pμ=L(p)μνkν.p^\mu = {L(p)^\mu}_\nu k^\nu.

Now apply a Lorentz transformation Λ\Lambda to a state with momentum pp. There are two equivalent ways to compare the internal labels:

  1. start at kk, use L(p)L(p) to reach pp, then use Λ\Lambda to reach Λp\Lambda p;
  2. start at kk and use the chosen standard transformation L(Λp)L(\Lambda p) to reach Λp\Lambda p directly.

The difference between these two procedures is

W(Λ,p)=L−1(Λp) Λ L(p).W(\Lambda,p)=L^{-1}(\Lambda p)\,\Lambda\,L(p).

By construction this transformation leaves the standard momentum fixed:

W(Λ,p)k=k.W(\Lambda,p)k=k.

This is the little group. It is not introduced by hand. It appears because once we factor out the purely kinematical change of momentum, the only remaining freedom is a Lorentz transformation that leaves the reference momentum unchanged. That remaining transformation is what acts on the internal labels σ\sigma.

This also explains why the boosts are not the main object of classification. We are not ignoring boosts. A boost is essential because it moves a particle from one momentum to another momentum on the same mass shell. But this motion is universal: every particle with the same mass has its momentum moved in the same way. By itself it does not tell us whether the particle is spin 0, spin 1/21/2, spin 1, or something else.

The intrinsic information is what remains after this momentum-changing part has been removed. In the formula above, L(p)L(p) and L(Λp)L(\Lambda p) account for the choice of boosts or standard Lorentz transformations that carry the reference momentum kk to the actual momentum. The leftover transformation W(Λ,p)W(\Lambda,p) keeps kk fixed, so it cannot be changing the momentum anymore. It acts only on the internal labels. That is why the little group, rather than the boosts themselves, classifies the spin or helicity content of the particle.

So Wigner’s classification of one-particle states reduces to this question:

For a chosen standard momentum kk, what are the unitary irreducible representations of the subgroup of Lorentz transformations that leaves kk fixed?

For massive particles this subgroup is SO(3)SO(3), which is why massive particles are classified by ordinary spin. For massless particles the subgroup is different, and that difference is exactly where the gauge transformation will come from.

Step 1: choose the standard null momentum

For a massless particle Weinberg chooses a standard four-momentum

kμ=(κ,0,0,κ),k^\mu = (\kappa,0,0,\kappa),

with κ>0\kappa > 0 fixed.

The little group consists of all Lorentz transformations WW that leave this momentum invariant:

Wμνkν=kμ.W^\mu{}_\nu k^\nu = k^\mu.

For a massless particle this little group is isomorphic to

ISO(2),ISO(2),

the Euclidean group in two dimensions: one rotation plus two translations.

This is already the essential difference with the massive case, where the little group is SO(3)SO(3).

Step 2: write the little-group elements explicitly

It is convenient to separate the little group into:

  • a rotation around the zz axis,
  • and a two-parameter “translation” part.

The rotation is

R(θ)=(10000cos⁡θ−sin⁡θ00sin⁡θcos⁡θ00001).R(\theta)= \begin{pmatrix} 1 & 0 & 0 & 0 \\ 0 & \cos\theta & -\sin\theta & 0 \\ 0 & \sin\theta & \cos\theta & 0 \\ 0 & 0 & 0 & 1 \end{pmatrix}.

One checks immediately that

R(θ)k=k.R(\theta)k = k.

The translation part may be written as

S(α,β)=(1+α2+β22αβ−α2+β22α10−αβ01−βα2+β22αβ1−α2+β22).S(\alpha,\beta)= \begin{pmatrix} 1+\frac{\alpha^2+\beta^2}{2} & \alpha & \beta & -\frac{\alpha^2+\beta^2}{2} \\ \alpha & 1 & 0 & -\alpha \\ \beta & 0 & 1 & -\beta \\ \frac{\alpha^2+\beta^2}{2} & \alpha & \beta & 1-\frac{\alpha^2+\beta^2}{2} \end{pmatrix}.

Again one verifies directly that

S(α,β)k=k.S(\alpha,\beta)k = k.

So a general little-group element may be built from these objects. The important point for us is not the exact group multiplication law, but the fact that the little group contains the non-compact translation sector (α,β)(\alpha,\beta).

Step 3: arbitrary unitary irreducible representations of the massless little group

At this point it is important to be precise. What we classify in Wigner’s construction is not an arbitrary unitary representation of the full Lorentz group, but an arbitrary unitary irreducible representation of the massless little group.

The Lie algebra of the little group is generated by

J3,N1,N2,J_3, \qquad N_1,\qquad N_2,

with commutation relations

[J3,N1]=iN2,[J3,N2]=−iN1,[N1,N2]=0.[J_3,N_1]=iN_2,\qquad [J_3,N_2]=-iN_1,\qquad [N_1,N_2]=0.

This is exactly the Lie algebra of ISO(2)ISO(2).

Now suppose we have a unitary representation. Then the generators are represented by self-adjoint operators, so in particular N1N_1 and N2N_2 are self-adjoint. Their eigenvalues must therefore be real. Also, since

[N1,N2]=0,[N_1,N_2]=0,

we can diagonalize them simultaneously, at least in the generalized sense appropriate for operators with continuous spectrum.

So let us first write a simultaneous generalized eigenstate as ∣n1,n2⟩|n_1,n_2\rangle:

N1∣n1,n2⟩=n1∣n1,n2⟩,N2∣n1,n2⟩=n2∣n1,n2⟩.N_1 |n_1,n_2\rangle = n_1 |n_1,n_2\rangle, \qquad N_2 |n_1,n_2\rangle = n_2 |n_1,n_2\rangle.

Here n1n_1 and n2n_2 are just two real numbers. The notation ρcos⁡ϕ\rho\cos\phi and ρsin⁡ϕ\rho\sin\phi is only a change to polar coordinates in this two-dimensional eigenvalue plane:

n1=ρcos⁡ϕ,n2=ρsin⁡ϕ,ρ=n12+n22≥0.n_1=\rho\cos\phi, \qquad n_2=\rho\sin\phi, \qquad \rho=\sqrt{n_1^2+n_2^2}\ge 0.

With this notation we call the same eigenstate ∣ρ,ϕ⟩|\rho,\phi\rangle, and the eigenvalue equations become

N1∣ρ,ϕ⟩=ρcos⁡ϕ ∣ρ,ϕ⟩,N2∣ρ,ϕ⟩=ρsin⁡ϕ ∣ρ,ϕ⟩,N_1 |\rho,\phi\rangle = \rho\cos\phi\, |\rho,\phi\rangle, \qquad N_2 |\rho,\phi\rangle = \rho\sin\phi\, |\rho,\phi\rangle,

with ρ≥0\rho\ge 0 by definition.

The number

ρ2=N12+N22\rho^2 = N_1^2 + N_2^2

is invariant under the rotation generated by J3J_3, because J3J_3 rotates the pair (N1,N2)(N_1,N_2) without changing its length. This is why ρ\rho labels the orbit of translation eigenvalues inside the little-group representation. The angle ϕ\phi tells us where we are on that orbit.

Now let us see how rotations act on these eigenstates. Since

U(R(θ))N1U(R(θ))−1=N1cos⁡θ+N2sin⁡θ,U(R(\theta))N_1U(R(\theta))^{-1}=N_1\cos\theta+N_2\sin\theta,

and

U(R(θ))N2U(R(θ))−1=−N1sin⁡θ+N2cos⁡θ,U(R(\theta))N_2U(R(\theta))^{-1}=-N_1\sin\theta+N_2\cos\theta,

it follows that

U(R(θ))∣ρ,ϕ⟩=∣ρ,ϕ+θ⟩U(R(\theta))|\rho,\phi\rangle = |\rho,\phi+\theta\rangle

up to an irrelevant overall phase convention.

On the other hand, the translation subgroup is generated by N1N_1 and N2N_2, so

U(S(α,β))=e iαN1+iβN2.U(S(\alpha,\beta)) = e^{\,i\alpha N_1 + i\beta N_2}.

Acting on the eigenstate ∣ρ,ϕ⟩|\rho,\phi\rangle gives

U(S(α,β))∣ρ,ϕ⟩=e iρ(αcos⁡ϕ+βsin⁡ϕ)∣ρ,ϕ⟩.U(S(\alpha,\beta))|\rho,\phi\rangle = e^{\,i\rho(\alpha\cos\phi+\beta\sin\phi)}|\rho,\phi\rangle.

So an arbitrary unitary irreducible representation of the massless little group is characterized by a non-negative number ρ\rho:

  • if ρ>0\rho>0, the rotation moves us continuously around the circle of angles ϕ\phi, and the translation subgroup acts non-trivially by phases. This is the continuous-spin case;
  • if ρ=0\rho=0, then
N1=N2=0N_1 = N_2 = 0

throughout the representation, so the translation subgroup acts trivially.

This second case is the one relevant for ordinary photons and, more generally, for the familiar massless particles of fixed helicity.

When ρ=0\rho=0, the little group reduces effectively to the rotation subgroup SO(2)SO(2). Its unitary irreducible representations are one-dimensional, so for a one-particle helicity state with standard momentum kk we have

U(W)∣k,σ⟩=eiσθ(W)∣k,σ⟩,U(W)|k,\sigma\rangle = e^{i\sigma \theta(W)} |k,\sigma\rangle,

and

U(S(α,β))∣k,σ⟩=∣k,σ⟩.U(S(\alpha,\beta))|k,\sigma\rangle = |k,\sigma\rangle.

This is the precise sense in which:

  • the rotation part acts by a phase,
  • the translation part acts trivially.

It is not a consequence of unitarity alone. It is the consequence of taking the ρ=0\rho=0 unitary irreducible representation of the massless little group, i.e. the discrete-helicity case rather than the continuous-spin case.

If parity is also imposed, then a state of helicity σ\sigma must be accompanied by a state of helicity −σ-\sigma. But again, parity is not what forces the translation part to be trivial.

This is crucial. The Hilbert space of a discrete-helicity massless particle only remembers helicity. But a Lorentz four-vector field will remember more structure than that.

Step 4: choose polarization vectors at the standard momentum

Take the usual transverse polarization vectors at kμ=(κ,0,0,κ)k^\mu=(\kappa,0,0,\kappa):

ϵ+μ(k)=12(0,1,i,0),ϵ−μ(k)=12(0,1,−i,0).\epsilon_+^\mu(k)=\frac{1}{\sqrt{2}}(0,1,i,0), \qquad \epsilon_-^\mu(k)=\frac{1}{\sqrt{2}}(0,1,-i,0).

They satisfy

kμϵ±μ(k)=0.k_\mu \epsilon_\pm^\mu(k)=0.

Now let us see how they transform under the little group.

Rotation part

Under R(θ)R(\theta) one finds

R(θ)ϵ+μ(k)=e−iθϵ+μ(k),R(θ)ϵ−μ(k)=e+iθϵ−μ(k).R(\theta)\epsilon_+^\mu(k)=e^{-i\theta}\epsilon_+^\mu(k), \qquad R(\theta)\epsilon_-^\mu(k)=e^{+i\theta}\epsilon_-^\mu(k).

So far, so good: this is exactly the helicity behavior we expect.

Translation part

Now comes the important calculation. Act with S(α,β)S(\alpha,\beta) on ϵ+μ(k)\epsilon_+^\mu(k):

S(α,β)ϵ+(k)=12(α+iβ1iα+iβ).S(\alpha,\beta)\epsilon_+(k) = \frac{1}{\sqrt{2}} \begin{pmatrix} \alpha+i\beta \\ 1 \\ i \\ \alpha+i\beta \end{pmatrix}.

Rewrite this as

S(α,β)ϵ+μ(k)=ϵ+μ(k)+α+iβ2 κkμ.S(\alpha,\beta)\epsilon_+^\mu(k) = \epsilon_+^\mu(k) + \frac{\alpha+i\beta}{\sqrt{2}\,\kappa} k^\mu.

Similarly,

S(α,β)ϵ−μ(k)=ϵ−μ(k)+α−iβ2 κkμ.S(\alpha,\beta)\epsilon_-^\mu(k) = \epsilon_-^\mu(k) + \frac{\alpha-i\beta}{\sqrt{2}\,\kappa} k^\mu.

This is the central fact.

The polarization vectors do not furnish a true two-dimensional representation of the little group. The translation part of the little group shifts them by something proportional to the null momentum kμk^\mu.

So already at the standard momentum we see the structure

ϵμ→ϵμ+c kμ.\epsilon^\mu \to \epsilon^\mu + c\, k^\mu.

That is the seed of the gauge transformation.

Step 5: go from the standard momentum to a generic momentum

Now take any null momentum

pμ=(∣p∣,p).p^\mu = (|\mathbf{p}|,\mathbf{p}).

Choose a Lorentz transformation L(p)L(p) such that

L(p)k=p.L(p)k = p.

One convenient choice is:

  1. first boost along the zz direction,
  2. then rotate the zz axis into the direction p^\hat{\mathbf p}.

If p\mathbf p has spherical angles (θ,ϕ)(\theta,\phi) and magnitude ∣p∣|\mathbf p|, define ξ\xi by

eξ=∣p∣κ.e^\xi = \frac{|\mathbf p|}{\kappa}.

Then

Bz(ξ)=(cosh⁡ξ00sinh⁡ξ01000010sinh⁡ξ00cosh⁡ξ)B_z(\xi)= \begin{pmatrix} \cosh\xi & 0 & 0 & \sinh\xi \\ 0 & 1 & 0 & 0 \\ 0 & 0 & 1 & 0 \\ \sinh\xi & 0 & 0 & \cosh\xi \end{pmatrix}

satisfies

Bz(ξ)k=(∣p∣,0,0,∣p∣).B_z(\xi)k = (|\mathbf p|,0,0,|\mathbf p|).

Now let

R(p^)=Rz(ϕ)Ry(θ),R(\hat{\mathbf p}) = R_z(\phi)R_y(\theta),

so that R(p^)R(\hat{\mathbf p}) sends the zz axis into the direction p^\hat{\mathbf p}. Then we may take

L(p)=R(p^)Bz(ξ),L(p)=R(\hat{\mathbf p})B_z(\xi),

and indeed

L(p)k=p.L(p)k = p.

The polarization vectors for momentum pp are then defined by

ϵ±μ(p)=L(p)μν ϵ±ν(k).\epsilon_\pm^\mu(p)=L(p)^\mu{}_\nu \, \epsilon_\pm^\nu(k).

Step 6: the Weinberg little-group element W(Λ,p)W(\Lambda,p)

Given a general Lorentz transformation Λ\Lambda, Weinberg defines

W(Λ,p)=L−1(Λp) Λ L(p).W(\Lambda,p)=L^{-1}(\Lambda p)\,\Lambda\,L(p).

This is one of the most important formulas in the whole argument.

Why?

Because it leaves kk invariant:

W(Λ,p)k=L−1(Λp)ΛL(p)k=L−1(Λp)Λp=k.W(\Lambda,p)k = L^{-1}(\Lambda p)\Lambda L(p)k = L^{-1}(\Lambda p)\Lambda p = k.

So W(Λ,p)W(\Lambda,p) belongs to the little group of the standard momentum.

This means that all the complicated Lorentz transformation properties at generic momentum are encoded in a little-group element acting at the standard momentum.

Step 7: transform the polarization vectors

Now compute:

Λμνϵ±ν(p)=ΛμνL(p)νρϵ±ρ(k).\Lambda^\mu{}_\nu \epsilon_\pm^\nu(p) = \Lambda^\mu{}_\nu L(p)^\nu{}_\rho \epsilon_\pm^\rho(k).

Insert the identity in the form L(Λp)L−1(Λp)L(\Lambda p)L^{-1}(\Lambda p):

Λϵ±(p)=L(Λp) W(Λ,p) ϵ±(k).\Lambda \epsilon_\pm(p) = L(\Lambda p)\,W(\Lambda,p)\,\epsilon_\pm(k).

Now use the explicit action of the little group on ϵ±(k)\epsilon_\pm(k):

W(Λ,p)ϵ±μ(k)=e∓iθ(Λ,p)ϵ±μ(k)+c±(Λ,p) kμ,W(\Lambda,p)\epsilon_\pm^\mu(k) = e^{\mp i\theta(\Lambda,p)}\epsilon_\pm^\mu(k) + c_\pm(\Lambda,p)\,k^\mu,

where c±(Λ,p)c_\pm(\Lambda,p) comes from the translation part of the little group.

Applying L(Λp)L(\Lambda p) gives

Λϵ±(p)=e∓iθ(Λ,p) ϵ±(Λp)+c±(Λ,p) L(Λp)k.\Lambda \epsilon_\pm(p) = e^{\mp i\theta(\Lambda,p)}\, \epsilon_\pm(\Lambda p) + c_\pm(\Lambda,p)\, L(\Lambda p)k.

But L(Λp)k=ΛpL(\Lambda p)k = \Lambda p, so finally

Λμνϵ±ν(p)=e∓iθ(Λ,p) ϵ±μ(Λp)+c±(Λ,p) (Λp)μ.\Lambda^\mu{}_\nu \epsilon_\pm^\nu(p) = e^{\mp i\theta(\Lambda,p)}\, \epsilon_\pm^\mu(\Lambda p) + c_\pm(\Lambda,p)\, (\Lambda p)^\mu.

This is the precise statement we wanted.

The polarization vector transforms:

  • as the expected helicity object,
  • plus an extra term proportional to the momentum.

And now we see exactly where that term comes from: from the translation part of the massless little group.

Step 8: why the extra term is harmless physically

Suppose the vector field couples to a conserved current JμJ^\mu. Then the physical amplitude contains

Jμϵ±μ(p).J_\mu \epsilon_\pm^\mu(p).

If we shift the polarization by a multiple of the momentum,

ϵ±μ(p)→ϵ±μ(p)+c pμ,\epsilon_\pm^\mu(p)\to \epsilon_\pm^\mu(p)+c\, p^\mu,

then the amplitude changes by

c Jμpμ.c\, J_\mu p^\mu.

But current conservation says

pμJμ=0.p_\mu J^\mu = 0.

So the shift by pμp^\mu does not change the physical amplitude.

Therefore the physically relevant polarization is really an equivalence class

ϵμ∼ϵμ+c pμ.\epsilon^\mu \sim \epsilon^\mu + c\, p^\mu.

Step 9: translate this into position space

Now expand a field operator in modes:

Aμ(x)∼∑σ∫dΠp[ϵμ(p,σ)a(p,σ)e−ipx+ϵμ∗(p,σ)a†(p,σ)eipx].A^\mu(x)\sim \sum_\sigma \int d\Pi_p \left[ \epsilon^\mu(p,\sigma)a(p,\sigma)e^{-ipx} + \epsilon^{\mu *}(p,\sigma)a^\dagger(p,\sigma)e^{ipx} \right].

Since in momentum space the polarization picks up an extra term proportional to pμp^\mu, in position space the field picks up an extra derivative:

pμ↔i∂μ.p^\mu \leftrightarrow i\partial^\mu.

Hence under Lorentz transformations the field must transform as

U(Λ)Aμ(x)U−1(Λ)=ΛμνAν(Λx)+∂μΩ(x,Λ),U(\Lambda) A^\mu(x) U^{-1}(\Lambda) = {\Lambda^\mu}_\nu A^\nu(\Lambda x) + \partial^\mu \Omega(x,\Lambda),

or equivalently, depending on conventions,

Aμ(x)→ΛμνAν(Λ−1x)+∂μΩΛ(x).A^\mu(x)\to {\Lambda^\mu}_\nu A^\nu(\Lambda^{-1}x)+\partial^\mu \Omega_\Lambda(x).

That last term is exactly the gauge transformation.

Why this is unavoidable

Let me summarize the logic in one chain:

  1. A Lorentz-covariant local field AμA_\mu has four components.
  2. A massless helicity-1 particle has fewer physical degrees of freedom: one helicity if parity is not imposed, or two helicities ±1\pm 1 if parity is imposed.
  3. The little group of a massless particle is ISO(2)ISO(2), not just SO(2)SO(2).
  4. The translation part of ISO(2)ISO(2) acts trivially on physical states but non-trivially on polarization vectors.
  5. Explicitly, it shifts ϵμ\epsilon^\mu by a multiple of the null momentum.
  6. Therefore a Lorentz-covariant vector field cannot transform as a strict four-vector on the physical Hilbert space.
  7. The mismatch is resolved precisely by gauge redundancy.

So gauge symmetry here is not just an aesthetic principle. It is the mechanism that allows a manifestly Lorentz-covariant field to describe the correct unitary massless representation.

Final remark

This is also why the field strength

Fμν=∂μAν−∂νAμF_{\mu\nu}=\partial_\mu A_\nu-\partial_\nu A_\mu

is often more directly physical than AμA_\mu itself. The gauge-variant part of AμA_\mu is exactly the redundant part required by Lorentz covariance.

So the correct final statement is:

massless helicity-1+Lorentz covariance+unitarity⟹Aμ transforms up to ∂μΩ.\text{massless helicity-1} + \text{Lorentz covariance} + \text{unitarity} \Longrightarrow A_\mu \text{ transforms up to } \partial_\mu \Omega.

If parity is also a symmetry, then the physical spectrum contains both helicities +1+1 and −1-1. But parity is not what forces the gauge term. The gauge term is already forced by the little-group structure of a massless particle.

References

  • Steven Weinberg, The Quantum Theory of Fields, Volume I, Chapter 5.
  • Mark Srednicki, Quantum Field Theory, sections on the photon and the little group.
  • Matthew D. Schwartz, Quantum Field Theory and the Standard Model, chapters on massless spin-1 fields.