Prime number theorem with the de la Vallée Poussin error term

We prove the classical quantitative prime number theorem: there is an absolute constant c>0 for which

\displaystyle \psi(x):=\sum_{n\le x}\Lambda(n)=x+O(x\exp(-c\sqrt{\log x})).

Consequently,

\displaystyle \vartheta(x):=\sum_{p\le x}\log p=x+O(x\exp(-c\sqrt{\log x})),

and

\displaystyle \pi(x)=\text{Li}(x)+O(\frac{x}{\log x}\exp(-c\sqrt{\log x})).

The main issue is not to prove merely that \psi(x)\sim x , but to understand why the error has the particular shape \exp(-c\sqrt{\log x}) . The answer is a balance between two costs in a contour shift. A zero-free region lets us move a Perron contour to the left of the line \text{Re}(s)=1 , gaining a factor from the power x^s . But the shift is only available up to a finite height T , and truncating the contour costs a factor involving T^{-1} . The de la Vallée Poussin zero-free region permits a horizontal displacement of order 1/\log T . Thus the saving is roughly \exp(-c\log x/\log T) , while the truncation cost is roughly \exp(-\log T) . Balancing these forces \log T\asymp\sqrt{\log x} , and that is where the square root comes from.

The proof has four linked stages. The Euler product turns primes into the Dirichlet coefficients of -\zeta'/\zeta . The completed zeta function then turns the nontrivial zeros of \zeta into a positive potential. The elementary inequality 3+4\cos\theta+\cos 2\theta\ge0 uses the positivity of the von Mangoldt coefficients to prevent this potential from placing a zero too close to \text{Re}(s)=1 . Finally, a smoothed Perron formula converts the zero-free region back into an estimate for primes. All implied constants below are absolute.

Primes and Riemann Zeta function

Define

\displaystyle \Lambda(n):=\begin{cases}\log p,&n=p^k\text{ for some prime }p\text{ and some integer }k\ge1,\\ 0,&\text{otherwise}.\end{cases}

The function \psi(x)=\sum_{n\le x}\Lambda(n) counts prime powers, weighting p^k by \log p . This is exactly the weight forced by logarithmic differentiation of the Euler product. For \text{Re}(s)>1 ,

\displaystyle \zeta(s)=\prod_p(1-p^{-s})^{-1},\qquad -\frac{\zeta'}{\zeta}(s)=\sum_{n=1}^{\infty}\frac{\Lambda(n)}{n^s}.

Indeed, differentiating the logarithm of the Euler product gives \sum_p(\log p)/(p^s-1) . Expanding (p^s-1)^{-1} as the geometric series \sum_{k\ge1}p^{-ks} produces one term (\log p)p^{-ks} for every prime power p^k . Thus the problem of estimating primes is transformed into the problem of estimating the logarithmic derivative of \zeta(s) near the boundary of its initial half-plane of convergence.

We shall also need the elementary continuation of \zeta to \text{Re}(s)>0 . Let A(u)=\lfloor u\rfloor . Summation by parts gives, initially for \text{Re}(s)>1 ,

\displaystyle \zeta(s)=s\int_1^{\infty}A(u)u^{-s-1}du=\frac{s}{s-1}-s\int_1^{\infty}{u}u^{-s-1}du.

The second integral converges already when \text{Re}(s)>0 , since 0\le{u}<1 . Hence \zeta continues meromorphically to this larger half-plane, with a unique singularity there: a simple pole at s=1 of residue 1 . In particular,

\displaystyle -\frac{\zeta'}{\zeta}(s)=\frac{1}{s-1}+O(1)

near s=1 . This pole is the analytic origin of the main term x .

The nontrivial zeros of \zeta are most useful after adjoining the Gamma factor that arises from Poisson summation. Let

\displaystyle \theta(u):=\sum_{n\in\mathbb Z}\exp(-\pi n^2u),\qquad u>0.

The Fourier transform of a Gaussian is another Gaussian, and Poisson summation gives the exact relation

\displaystyle \theta(u)=u^{-1/2}\theta(1/u).

For \text{Re}(s)>1 , Mellin integration term by term gives

\displaystyle \pi^{-s/2}\Gamma\left(\frac{s}{2}\right)\zeta(s)=\frac12\int_0^{\infty}(\theta(u)-1)u^{s/2-1}du.

Split this integral at u=1 . On the interval 0<u<1 , insert the theta transformation and then substitute u=1/v . One obtains

\displaystyle \pi^{-s/2}\Gamma\left(\frac{s}{2}\right)\zeta(s)=\frac{1}{s-1}-\frac{1}{s}+F(s),

where

\displaystyle F(s):=\frac12\int_1^{\infty}(\theta(u)-1)\left(u^{s/2-1}+u^{-(s+1)/2}\right)du.

The function \theta(u)-1 decays exponentially as u\to\infty , so this integral converges locally uniformly for every complex s . Hence F is entire. The displayed expression is also unchanged when s is replaced by 1-s . Therefore the completed zeta function

\displaystyle \xi(s):=\frac12s(s-1)\pi^{-s/2}\Gamma\left(\frac{s}{2}\right)\zeta(s)

is entire and satisfies the functional equation

\displaystyle \xi(s)=\xi(1-s).

The zeros of \xi are exactly the nontrivial zeros \rho=\beta+i\gamma of \zeta . The Euler product shows that \zeta(s)\ne0 for \text{Re}(s)>1 ; the functional equation reflects zeros across \text{Re}(s)=1/2 . Hence every nontrivial zero lies in the strip 0\le\beta\le1 .

We also need a mild estimate for the Gamma factor. The classical series

\displaystyle \frac{\Gamma'}{\Gamma}(z)=-\gamma+\sum_{m=0}^{\infty}\left(\frac{1}{m+1}-\frac{1}{m+z}\right)

shows that

\displaystyle \frac{\Gamma'}{\Gamma}(z)\ll\log(|z|+2),\qquad \text{Re}(z)\ge\frac12.

To see this directly, split the sum at m\asymp |z|+2 . The initial portion is a harmonic sum of size O(\log(|z|+2)) , while the remainder is bounded because each summand is O(|z|/m^2) .

The crucial gain from \xi is that its logarithmic derivative records the zeros with positive weights to the right of the critical strip. The required form follows from the Hadamard factorization of the entire function \xi . Its order is at most one; more concretely,

\displaystyle \log|\xi(s)|\ll |s|\log(|s|+3).

The canonical product therefore has genus one, and, with zeros counted with multiplicity,

\displaystyle \frac{\xi'}{\xi}(s)=B+\sum_{\rho}\left(\frac{1}{s-\rho}+\frac{1}{\rho}\right).

The sum is understood in the usual symmetric sense. Using the functional equation and the symmetry \rho\leftrightarrow1-\overline\rho of the zero set, the real part simplifies to

\displaystyle \text{Re}\frac{\xi'}{\xi}(\sigma+it)=\sum_{\rho}\frac{\sigma-\beta}{(\sigma-\beta)^2+(t-\gamma)^2},\qquad \sigma>1.

Every summand is nonnegative. One may picture a zero \rho as producing a positive peak centred at height \gamma ; the closer \sigma+it is to that zero, the larger its contribution.

Differentiating the definition of \xi now gives

\displaystyle \frac{\xi'}{\xi}(s)=\frac{1}{s}+\frac{1}{s-1}-\frac12\log\pi+\frac12\frac{\Gamma'}{\Gamma}\left(\frac{s}{2}\right)+\frac{\zeta'}{\zeta}(s).

Combining this with the Gamma estimate and the positive zero formula yields, uniformly for 1<\sigma\le2 ,

\displaystyle -\text{Re}\frac{\zeta'}{\zeta}(\sigma+it)=\frac{\sigma-1}{(\sigma-1)^2+t^2}-\sum_{\rho}\frac{\sigma-\beta}{(\sigma-\beta)^2+(t-\gamma)^2}+O(\log(|t|+3)).

The sign in front of the zero sum is the feature to remember. A zero lying close to \sigma+it makes -\zeta'/\zeta very negative. We will contradict such a large negative contribution using a nonnegative trigonometric polynomial and the nonnegative coefficients \Lambda(n) .

Trigonometric Identity and Zero Free region

The whole zero-free argument rests on the elementary identity

\displaystyle 3+4\cos\theta+\cos 2\theta=2(1+\cos\theta)^2\ge0.

For \sigma>1 , the Dirichlet series for -\zeta'/\zeta gives

\displaystyle \begin{aligned} &-3\frac{\zeta'}{\zeta}(\sigma)-4\text{Re}\frac{\zeta'}{\zeta}(\sigma+it)-\text{Re}\frac{\zeta'}{\zeta}(\sigma+2it)\\ &=\sum_{n=1}^{\infty}\frac{\Lambda(n)}{n^{\sigma}}\left(3+4\cos(t\log n)+\cos(2t\log n)\right)\ge0. \end{aligned}

This is not a formal trick. It compares the logarithmic derivative at three heights, 0,t,2t . If there were a zero close to 1+it , the middle term would become strongly negative. The coefficient 4 in front of that term is deliberately larger than the coefficient 3 in front of the pole at s=1 . That small numerical advantage is what leaves a positive gap between zeros and the line \text{Re}(s)=1 .

Let \rho_0=\beta_0+i\gamma_0 be a zero, write \ell:=\log(|\gamma_0|+3) , and put \sigma=1+\delta . The pole estimate at 1+\delta gives -\zeta'/\zeta(1+\delta)\le\delta^{-1}+O(1) . At the point 1+\delta+i\gamma_0 , the single zero \rho_0 contributes -1/(\delta+1-\beta_0) , and every other zero contributes with the same nonpositive sign. At height 2\gamma_0 , we may simply discard the nonpositive zero contributions. Substitution into the positive trigonometric inequality gives

\displaystyle 0\le\frac{3}{\delta}-\frac{4}{\delta+1-\beta_0}+C\ell.

Choose \delta=(2C\ell)^{-1} , enlarging C once if necessary. Then 3/\delta+C\ell=7C\ell , so

\displaystyle \delta+1-\beta_0\ge\frac{4}{7C\ell}.

Subtracting \delta leaves

\displaystyle 1-\beta_0\ge\frac{1}{14C\log(|\gamma_0|+3)}.

Thus there is an absolute constant a>0 for which

\displaystyle \zeta(s)\ne0\quad\text{whenever}\quad \text{Re}(s)\ge1-\frac{a}{\log(|\text{Im}(s)|+3)}.

This is the de la Vallée Poussin zero-free region. It says that a zero at height T must stay at least on the order of 1/\log T away from the line \text{Re}(s)=1 . The reciprocal logarithm may look weak, but it is already strong enough to give an exponentially decaying error in the prime number theorem.

Before shifting a contour, we need a uniform bound for \zeta'/\zeta in the zero-free region. The positive potential formula first gives a local zero count. At s=2+it ,

\displaystyle \text{Re}\frac{\xi'}{\xi}(2+it)=\sum_{\rho}\frac{2-\beta}{(2-\beta)^2+(t-\gamma)^2}.

Every zero with |\gamma-t|\le1 contributes at least 1/5 , because 1\le2-\beta\le2 . On the other hand, the explicit formula for \xi'/\xi , the Gamma estimate, and the absolute convergence of \zeta'/\zeta on \text{Re}(s)=2 show that the left side is O(\log(|t|+3)) . Hence

\displaystyle |\{\rho:|\gamma-t|\le1\}| \ll\log(|t|+3).

This tells us that only logarithmically many zero peaks can be near any given height. To convert this into a partial-fraction estimate, compare \xi'/\xi(s) with \xi'/\xi(2+it) , where s=\sigma+it and 1/2\le\sigma\le2 . For zeros with |\gamma-t|>1 ,

\displaystyle \left|\frac{1}{s-\rho}-\frac{1}{2+it-\rho}\right|\ll\frac{1}{|t-\gamma|^2}.

Group the zeros into strips m<|t-\gamma|\le m+1 . The local zero count bounds the number in each strip by O(\log(|t|+m+3)) , and summing the resulting convergent series gives

\displaystyle \frac{\zeta'}{\zeta}(s)=\sum_{|\gamma-t|\le1}\frac{1}{s-\rho}+O\left(\log(|t|+3)+\frac{1}{|s-1|}\right).

Now fix T\ge3 and put

\displaystyle \sigma_0:=1-\frac{a}{2\log T}.

If |t|\le T and \sigma_0\le\sigma\le2 , then the zero-free region keeps every zero with |\gamma-t|\le1 at horizontal distance \gg1/\log T from s . There are only O(\log T) such zeros, so their total contribution is O((\log T)^2) . The remaining terms are smaller. Consequently,

\displaystyle \frac{\zeta'}{\zeta}(\sigma+it)\ll(\log T)^2,\quad \sigma_0\le\sigma\le2,\quad |t|\le T.

This is the estimate that makes the contour shift quantitative.

Perron’s Formula

A direct Perron formula for \psi(x) has a kernel 1/s , which decays only like 1/|t| and makes truncation awkward. We therefore integrate once and work instead with

\displaystyle \psi_1(x):=\sum_{n\le x}(x-n)\Lambda(n).

The new kernel is 1/(s(s+1)) , which decays like 1/t^2 . The elementary Mellin calculation behind the formula is

\displaystyle \frac{1}{2\pi i}\int_{c-i\infty}^{c+i\infty}\frac{y^s}{s(s+1)},ds=\begin{cases}1-y^{-1},& y>1, \\ 0, & 0<y<1 \end{cases}

For y>1 , shift the contour to the left and collect the residues at s=0 and s=-1 ; for 0<y<1 , shift it to the right, where there are no poles. Insert the Dirichlet series for -\zeta'/\zeta and apply this kernel with y=x/n . The result is

\displaystyle \psi_1(x)=\frac{1}{2\pi i}\int_{c-i\infty}^{c+i\infty}-\frac{\zeta'}{\zeta}(s)\frac{x^{s+1}}{s(s+1)}ds,\qquad c>1.

The pole of -\zeta'/\zeta at s=1 has residue 1 . Hence, when the contour is moved to the left, it contributes the main term x^2/2 .

Let L:=\log x , take c=1+1/L , and choose \displaystyle \log T=\sqrt{\frac{aL}{2}},\quad \sigma_0=1-\frac{a}{2\log T}.

For bounded x , the desired estimate can be absorbed into the constant, so we may assume that T is large. Shift the contour through the rectangle with vertices c\pm iT and \sigma_0\pm iT . The zero-free region shows that no zero of \zeta is crossed. The only singularity crossed is the pole at s=1 , so

\displaystyle \psi_1(x)=\frac{x^2}{2}+E_{\text{vert}}+E_{\text{hor}}+E_{\text{tail}}.

On the new vertical side, the logarithmic derivative estimate and the integrability of (1+t^2)^{-1} give

\displaystyle E_{\text{vert}}\ll x^{\sigma_0+1}(\log T)^2=x^2(\log T)^2\exp\left(-\frac{a\log x}{2\log T}\right).

On the two horizontal sides, |s(s+1)|\gg T^2 , so \displaystyle E_{\text{hor}}\ll \frac{x^2(\log T)^2}{T^2}.

Finally, on the original line \text{Re}(s)=c , the absolutely convergent Dirichlet series gives |\zeta'/\zeta(c+it)|\ll(c-1)^{-1}\ll L . The omitted tails above height T therefore contribute \displaystyle E_{\text{tail}}\ll\frac{x^2L}{T}.

With our choice of T , both a\log x/(2\log T) and \log T equal \sqrt{a\log x/2} . The logarithmic factors can be absorbed by slightly weakening the constant in the exponential. Thus, for some absolute b>0 ,

\displaystyle \psi_1(x)=\frac{x^2}{2}+O\left(x^2\exp\left(-b\sqrt{\log x}\right)\right).

Removing Smoothing

The final step uses only positivity. Let 0<h\le x/2 . Every term with n\le x-h contributes exactly h\Lambda(n) to \psi_1(x)-\psi_1(x-h) , while terms with x-h<n\le x make an additional nonnegative contribution. The same observation on the interval to the right of x gives

\displaystyle \frac{\psi_1(x)-\psi_1(x-h)}{h}\le\psi(x)\le\frac{\psi_1(x+h)-\psi_1(x)}{h}.

Insert the asymptotic formula for \psi_1 . The main terms on the right and left are respectively x+h/2 and x-h/2 . Choose

\displaystyle h=x\exp\left(-\frac{b}{2}\sqrt{\log x}\right).

The smoothing error, after division by h , has the same order as h . Hence, after decreasing the positive constant once more,

\displaystyle \psi(x)=x+O\left(x\exp\left(-c\sqrt{\log x}\right)\right).

This is the de la Vallée Poussin form of the prime number theorem for the Chebyshev function.

The smoothing is a convenience rather than a necessity. One can begin directly from the truncated sharp Perron integral

\displaystyle I(x,T):=\frac{1}{2\pi i}\int_{c-iT}^{c+iT}-\frac{\zeta'(s)}{\zeta(s)}\frac{x^s}{s},ds,\qquad c=1+\frac{1}{\log x}.

The infinite version of this integral inverts the discontinuous cutoff \displaystyle \mathbf 1_{n\le x} : it equals 1 when n<x , 0 when n>x , and 1/2 at n=x . At finite height T , however, the cutoff is blurred in the short window \displaystyle |n-x|\ll x/T. The standard truncated Perron estimate therefore gives

\displaystyle \psi(x)=I(x,T)+O\left(\frac{x(\log x)^2}{T}+\log x\right).

The term \displaystyle x(\log x)^2/T is the price of using a sharp cutoff: the contour cannot distinguish perfectly between integers just below and just above x .

The contour shift is otherwise the same. Move the segment to \displaystyle \text{Re}(s)=1-\frac{2}{a\log T}, where a>0 comes from the zero-free region. The pole at s=1 contributes x . Since the kernel is now only 1/s , rather than 1/(s(s+1)) , the vertical integral has one extra logarithmic loss and the horizontal segments have size about \displaystyle x(\log T)^2/T. One obtains

\displaystyle \psi(x)=x+O\left(x\exp\left(-\frac{2\log x}{a\log T}\right)(\log T)^3+\frac{x(\log x)^2}{T}\right).

Choosing \displaystyle \log T\asymp\sqrt{\log x} balances the two exponential savings and yields again

\displaystyle \psi(x)=x+O\left(x\exp(-c\sqrt{\log x})\right).

Thus smoothing does not improve the final de la Vallée Poussin error term. Its role is expository and technical: replacing the discontinuous cutoff by \displaystyle (x-n)_+ upgrades \displaystyle 1/s to \displaystyle 1/(s(s+1)), removes the boundary layer near n=x , and makes the contour estimates absolutely convergent.

Prime Number Theorem

The difference \psi(x)-\vartheta(x) counts only proper prime powers. Without using any prime number theorem, one has

\displaystyle 0\le\psi(x)-\vartheta(x)=\sum_{{p^k\le x, ~ k\ge2}}\log p\ll\sqrt{x}(\log x)^2.

Indeed, for each k\ge2 , there are at most x^{1/k} possible primes p , each weighted by at most \log x , and the sum over k is dominated by its first term up to a logarithmic factor. This is negligible compared with x\exp(-c\sqrt{\log x}) . Therefore

\displaystyle \vartheta(x)=x+O\left(x\exp\left(-c\sqrt{\log x}\right)\right).

Finally, partial summation gives

\displaystyle \pi(x)=\frac{\vartheta(x)}{\log x}+\int_2^x\frac{\vartheta(t)}{t(\log t)^2}dt.

Substituting the main term \vartheta(t)=t gives \text{Li}(x)+O(1) , because the derivative of x/\log x+\int_2^xdt/(\log t)^2 is 1/\log x . For the error term, split the integral at \sqrt{x} . The lower portion is far smaller than the claimed final error. On the interval [\sqrt{x},x] , one has \sqrt{\log t}\ge\sqrt{\log x}/\sqrt2 , so the exponential factor is uniformly bounded by a slightly weaker exponential in \sqrt{\log x} . Consequently, after one final reduction of c ,

\displaystyle \pi(x)=\text{Li}(x)+O\left(\frac{x}{\log x}\exp\left(-c\sqrt{\log x}\right)\right).

It is useful to see the proof as a single chain rather than as a collection of unrelated estimates. The Euler product says that weighted prime powers are the coefficients of -\zeta'/\zeta . Completion and factorization say that zeros of \zeta appear as a positive potential in \text{Re}(\xi'/\xi) . The polynomial 3+4\cos\theta+\cos2\theta converts the positivity of the coefficients \Lambda(n) into an inequality that cannot coexist with a zero extremely close to one. The zero-free region then permits a contour shift; the shifted contour produces a saving, and the finite truncation produces a competing loss. The two costs balance at height \log T\asymp\sqrt{\log x} , giving the classical error term. The square root is therefore not an incidental artifact of the proof. It is the numerical fingerprint of a reciprocal-logarithmic zero-free region. A stronger zero-free region would move the contour farther left and improve the error term; a weaker one would give less saving. The prime number theorem with de la Vallée Poussin’s error term is exactly what this classical zero-free region is designed to yield.

Leave a comment