We prove the classical quantitative prime number theorem: there is an absolute constant for which
Consequently,
and
The main issue is not to prove merely that , but to understand why the error has the particular shape
. The answer is a balance between two costs in a contour shift. A zero-free region lets us move a Perron contour to the left of the line
, gaining a factor from the power
. But the shift is only available up to a finite height
, and truncating the contour costs a factor involving
. The de la Vallée Poussin zero-free region permits a horizontal displacement of order
. Thus the saving is roughly
, while the truncation cost is roughly
. Balancing these forces
, and that is where the square root comes from.
The proof has four linked stages. The Euler product turns primes into the Dirichlet coefficients of . The completed zeta function then turns the nontrivial zeros of
into a positive potential. The elementary inequality
uses the positivity of the von Mangoldt coefficients to prevent this potential from placing a zero too close to
. Finally, a smoothed Perron formula converts the zero-free region back into an estimate for primes. All implied constants below are absolute.
Primes and Riemann Zeta function
Define
The function counts prime powers, weighting
by
. This is exactly the weight forced by logarithmic differentiation of the Euler product. For
,
Indeed, differentiating the logarithm of the Euler product gives . Expanding
as the geometric series
produces one term
for every prime power
. Thus the problem of estimating primes is transformed into the problem of estimating the logarithmic derivative of
near the boundary of its initial half-plane of convergence.
We shall also need the elementary continuation of to
. Let
. Summation by parts gives, initially for
,
The second integral converges already when , since
. Hence
continues meromorphically to this larger half-plane, with a unique singularity there: a simple pole at
of residue
. In particular,
near . This pole is the analytic origin of the main term
.
The nontrivial zeros of are most useful after adjoining the Gamma factor that arises from Poisson summation. Let
The Fourier transform of a Gaussian is another Gaussian, and Poisson summation gives the exact relation
For , Mellin integration term by term gives
Split this integral at . On the interval
, insert the theta transformation and then substitute
. One obtains
where
The function decays exponentially as
, so this integral converges locally uniformly for every complex
. Hence
is entire. The displayed expression is also unchanged when
is replaced by
. Therefore the completed zeta function
is entire and satisfies the functional equation
The zeros of are exactly the nontrivial zeros
of
. The Euler product shows that
for
; the functional equation reflects zeros across
. Hence every nontrivial zero lies in the strip
.
We also need a mild estimate for the Gamma factor. The classical series
shows that
To see this directly, split the sum at . The initial portion is a harmonic sum of size
, while the remainder is bounded because each summand is
.
The crucial gain from is that its logarithmic derivative records the zeros with positive weights to the right of the critical strip. The required form follows from the Hadamard factorization of the entire function
. Its order is at most one; more concretely,
The canonical product therefore has genus one, and, with zeros counted with multiplicity,
The sum is understood in the usual symmetric sense. Using the functional equation and the symmetry of the zero set, the real part simplifies to
Every summand is nonnegative. One may picture a zero as producing a positive peak centred at height
; the closer
is to that zero, the larger its contribution.
Differentiating the definition of now gives
Combining this with the Gamma estimate and the positive zero formula yields, uniformly for ,
The sign in front of the zero sum is the feature to remember. A zero lying close to makes
very negative. We will contradict such a large negative contribution using a nonnegative trigonometric polynomial and the nonnegative coefficients
.
Trigonometric Identity and Zero Free region
The whole zero-free argument rests on the elementary identity
For , the Dirichlet series for
gives
This is not a formal trick. It compares the logarithmic derivative at three heights, . If there were a zero close to
, the middle term would become strongly negative. The coefficient
in front of that term is deliberately larger than the coefficient
in front of the pole at
. That small numerical advantage is what leaves a positive gap between zeros and the line
.
Let be a zero, write
, and put
. The pole estimate at
gives
. At the point
, the single zero
contributes
, and every other zero contributes with the same nonpositive sign. At height
, we may simply discard the nonpositive zero contributions. Substitution into the positive trigonometric inequality gives
Choose , enlarging
once if necessary. Then
, so
Subtracting leaves
Thus there is an absolute constant for which
This is the de la Vallée Poussin zero-free region. It says that a zero at height must stay at least on the order of
away from the line
. The reciprocal logarithm may look weak, but it is already strong enough to give an exponentially decaying error in the prime number theorem.
Before shifting a contour, we need a uniform bound for in the zero-free region. The positive potential formula first gives a local zero count. At
,
Every zero with contributes at least
, because
. On the other hand, the explicit formula for
, the Gamma estimate, and the absolute convergence of
on
show that the left side is
. Hence
This tells us that only logarithmically many zero peaks can be near any given height. To convert this into a partial-fraction estimate, compare with
, where
and
. For zeros with
,
Group the zeros into strips . The local zero count bounds the number in each strip by
, and summing the resulting convergent series gives
Now fix and put
If and
, then the zero-free region keeps every zero with
at horizontal distance
from
. There are only
such zeros, so their total contribution is
. The remaining terms are smaller. Consequently,
This is the estimate that makes the contour shift quantitative.
Perron’s Formula
A direct Perron formula for has a kernel
, which decays only like
and makes truncation awkward. We therefore integrate once and work instead with
The new kernel is , which decays like
. The elementary Mellin calculation behind the formula is
For , shift the contour to the left and collect the residues at
and
; for
, shift it to the right, where there are no poles. Insert the Dirichlet series for
and apply this kernel with
. The result is
The pole of at
has residue
. Hence, when the contour is moved to the left, it contributes the main term
.
Let , take
, and choose
For bounded , the desired estimate can be absorbed into the constant, so we may assume that
is large. Shift the contour through the rectangle with vertices
and
. The zero-free region shows that no zero of
is crossed. The only singularity crossed is the pole at
, so
On the new vertical side, the logarithmic derivative estimate and the integrability of give
On the two horizontal sides, , so
Finally, on the original line , the absolutely convergent Dirichlet series gives
. The omitted tails above height
therefore contribute
With our choice of , both
and
equal
. The logarithmic factors can be absorbed by slightly weakening the constant in the exponential. Thus, for some absolute
,
Removing Smoothing
The final step uses only positivity. Let . Every term with
contributes exactly
to
, while terms with
make an additional nonnegative contribution. The same observation on the interval to the right of
gives
Insert the asymptotic formula for . The main terms on the right and left are respectively
and
. Choose
The smoothing error, after division by , has the same order as
. Hence, after decreasing the positive constant once more,
This is the de la Vallée Poussin form of the prime number theorem for the Chebyshev function.
The smoothing is a convenience rather than a necessity. One can begin directly from the truncated sharp Perron integral
The infinite version of this integral inverts the discontinuous cutoff : it equals
when
,
when
, and
at
. At finite height
, however, the cutoff is blurred in the short window
The standard truncated Perron estimate therefore gives
The term is the price of using a sharp cutoff: the contour cannot distinguish perfectly between integers just below and just above
.
The contour shift is otherwise the same. Move the segment to where
comes from the zero-free region. The pole at
contributes
. Since the kernel is now only
, rather than
, the vertical integral has one extra logarithmic loss and the horizontal segments have size about
One obtains
Choosing balances the two exponential savings and yields again
Thus smoothing does not improve the final de la Vallée Poussin error term. Its role is expository and technical: replacing the discontinuous cutoff by upgrades
to
removes the boundary layer near
, and makes the contour estimates absolutely convergent.
Prime Number Theorem
The difference counts only proper prime powers. Without using any prime number theorem, one has
Indeed, for each , there are at most
possible primes
, each weighted by at most
, and the sum over
is dominated by its first term up to a logarithmic factor. This is negligible compared with
. Therefore
Finally, partial summation gives
Substituting the main term gives
, because the derivative of
is
. For the error term, split the integral at
. The lower portion is far smaller than the claimed final error. On the interval
, one has
, so the exponential factor is uniformly bounded by a slightly weaker exponential in
. Consequently, after one final reduction of
,
It is useful to see the proof as a single chain rather than as a collection of unrelated estimates. The Euler product says that weighted prime powers are the coefficients of . Completion and factorization say that zeros of
appear as a positive potential in
. The polynomial
converts the positivity of the coefficients
into an inequality that cannot coexist with a zero extremely close to one. The zero-free region then permits a contour shift; the shifted contour produces a saving, and the finite truncation produces a competing loss. The two costs balance at height
, giving the classical error term. The square root is therefore not an incidental artifact of the proof. It is the numerical fingerprint of a reciprocal-logarithmic zero-free region. A stronger zero-free region would move the contour farther left and improve the error term; a weaker one would give less saving. The prime number theorem with de la Vallée Poussin’s error term is exactly what this classical zero-free region is designed to yield.