A first course in probability often gives us one very powerful mental picture: when many independent random variables are added, the answer should begin to look Gaussian. This is the central limit theorem. In its most familiar form, it says that if are independent and identically distributed, with mean
and variance
, then
This is one of the main organizing facts of probability. But the slogan “sums become Gaussian” is too crude. A sum becomes Gaussian only when we look at it on a scale where no single summand remains visible, while the total quadratic fluctuation remains visible. If that balance fails, other limits can appear. Some limits remember a finite number of rare events. Some remember jumps of many different sizes. Some have infinite variance. Some even have infinite mean. The point of this article is to understand, concretely and through calculation, how these possibilities arise.
The surprising part is that all of this can happen even when every finite-stage random variable is bounded. Suppose that for each we have independent random variables
and suppose that
for every
and every
. At no finite stage do we use an unbounded random variable. Nevertheless, the sums
can have Gaussian limits, Poisson limits, compound Poisson limits, truncated stable limits, full stable limits, and positive stable limits with infinite mean. The reason is that the distribution of is allowed to depend on
. This is a triangular array, not a single fixed distribution repeated forever.
The basic issue is scale. If we divide by a number , then the bounded variable
becomes
. Even though
, the normalized variable satisfies
and
. Thus the original variable is bounded, but the magnified variable is not bounded uniformly in
. The interval
becomes
, and this expands toward the whole real line. This is the basic mechanism by which bounded triangular arrays can produce infinite-variance or infinite-mean limits.
So the right question is not simply: are the summands bounded? The right question is: after normalization, what does one summand look like? Does every normalized summand become tiny? Does a finite number of jumps remain visible? Do jumps survive at many different scales? Does the maximum matter on the same scale as the whole sum? These are the questions that separate the Gaussian world from the Poisson world and the stable world.
A useful diagnostic throughout the article will be the scale-counting quantity
This has a direct meaning: among independent observations, it is approximately the expected number whose magnitude is larger than
. If this number is huge, then many observations are visible above scale
. If it tends to a finite nonzero number, then jumps at that scale survive individually. If it tends to zero for every fixed
, then perhaps we are looking too coarsely, and the interesting structure has moved to a smaller scale. This quantity often sees more than the variance sees.
The Fourier Idea
Before going through examples, let us record the main calculation principle. The characteristic function of a random variable is
If and the
are independent, then
In the identically distributed row case, this is
Taking logarithms gives
This is why characteristic functions are so useful for sums. Independence turns sums of random variables into products of characteristic functions, and logarithms turn products into sums. If one summand makes a small contribution to the characteristic function, then multiplying by tells us whether that contribution disappears, survives, or explodes.
In many triangular-array problems, for fixed one has
Then
provided the small term is small enough. So the central question becomes: what is the first non-negligible term in the one-step characteristic function, and what happens after multiplying by
? This is the Fourier-side version of the physical question: what remains visible after adding many independent pieces?
Fair coins
Start with the cleanest model. Let be independent fair signs, so
and put Since
and
, independence gives
Thus the natural size of
is
, not
. The positive and negative signs cancel enough that the typical imbalance is of order
. So we look at
This display already contains the Gaussian geometry. Each normalized summand has size , so
Thus no individual toss remains visible. But the total variance remains visible, since
So every single contribution disappears, but the total quadratic fluctuation remains order one. This is exactly the situation where Gaussian behavior should arise.
Now calculate the characteristic function. For one normalized sign,
By independence, Set
. Since
we get
This line is the whole story. The quadratic term survives because The quartic term disappears because
The sixth-order term disappears even faster. More generally, a term of order
contributes roughly
When
, this is order one. When
, it vanishes. The linear term would have size
, but it is absent because the variable is centered. Positive and negative signs cancel at first order.
Therefore
which is the characteristic function of . So Gaussian behavior comes from a specific balance. The normalization makes each individual summand invisible, but preserves the total variance. If we divided by a much larger scale than
, even the variance would disappear and the limit would collapse to zero. If we divided by a much smaller scale, the sum would spread out. The scale
is the scale where the quadratic term is neither killed nor blown up.
Poisson behavior
Now change just one feature. Instead of receiving a nonzero contribution on every trial, suppose most trials give zero and only rarely something happens. Let
Define This counts the number of successes among
trials. By linearity of expectation,
The quantity
is the expected number of visible events. In the fair-coin model, something happened on every trial, so the number of active contributions was exactly
. Here the number of active contributions is random, and its typical size is controlled by
.
There are three regimes.
If , then even one success is unlikely. Indeed,
When
is small,
. If
, then
, so
in probability. The events are too rare to survive.
If , then many successes occur. After subtracting the mean and dividing by the standard deviation, one is again in a Gaussian-type regime:
under the usual nondegeneracy assumptions. Many small centered errors accumulate, and no single trial decides the final normalized value.
The new regime is the middle one: Here each event becomes rarer and rarer, but the number of opportunities grows at just the right rate. We should expect a finite random number of visible events. This is the Poisson regime.
For fixed ,
If
, then
, and
behaves like
. For fixed
,
Thus
Also,
Therefore
So This is already a non-Gaussian limit from bounded variables taking only the values
and
. The reason is not mysterious. In the Gaussian coin model, many contributions survive but each becomes individually invisible after normalization. In the Poisson regime, only finitely many events survive, and they remain visible. The limit remembers the number of rare events.
Poisson jumps or Gaussian fluctuation
Now let the rare event have a sign. Define
Most of the time nothing happens. When a jump occurs, it is equally likely to be positive or negative. The variable is still bounded and simple: it takes only the values . Let
By symmetry,
. Also,
exactly when a nonzero jump occurs. Therefore
and hence Again,
measures the expected number of active jumps.
If , then only finitely many nonzero jumps survive. The number of positive jumps converges to
, and the number of negative jumps converges to an independent
. Thus
where and
are independent
random variables. This is a jump limit. The limit remembers the surviving positive and negative jumps.
The characteristic function gives the same result. For one summand,
Therefore
If , this tends to
This is the characteristic function of the difference of two independent variables.
If instead , then the number of nonzero jumps diverges. The natural scale is the standard deviation
. Under this scaling, one nonzero jump has size
Now individual jumps are no longer visible. We are back in the Gaussian geometry. Indeed,
Using with
, we get
Multiplying by , the main term becomes
, and the error vanishes if
. Thus
The same bounded variables produce two different limits. If stays finite, the limit is made of finitely many visible jumps. If
tends to infinity, the number of jumps grows, each normalized jump disappears, and the Gaussian mechanism returns.
Compound Poisson limits
The rare signed coin has only one nonzero magnitude: every jump has size . To get a richer jump limit, allow the size of the jump to be random. Let
where ,
, and all variables are independent. The Bernoulli variable decides whether a jump occurs. If
, then
. If
, then the jump size is
, uniformly distributed in
.
Assume . Then the number of jumps converges to a Poisson random variable
. Conditional on
, the limiting contribution is
Thus the limit has the form This is a compound Poisson random variable. The word “Poisson” records the random number of jumps. The word “compound” records that each jump carries a random size.
The characteristic function confirms this. For ,
For one summand,
Therefore
If , then
and hence
This is exactly the characteristic function of . If we condition on
, then
Averaging over gives the same exponential.
So when finitely many jumps survive, the limit remembers their sizes. This suggests the next question: how complicated can the surviving jump-size distribution become?
A first attempt is to change the width of a familiar distribution. Suppose This changes the physical size of the summand, but it does not create a hierarchy of scales. A typical observation is of size comparable to
. There are not rare jumps at size
, more common jumps at size
, and even more common jumps at
. There is one scale.
Indeed, The standard normalized variable is
The factor disappears after scaling. The characteristic function is
Since we get
Thus Changing the width of a uniform distribution changes the unit of measurement, not the limiting mechanism. It does not produce a new jump geometry.
Now try a slightly richer model. Suppose a rare jump can have one of two magnitudes, say or
. Let a jump occur with probability
. Conditional on a jump, suppose it has signed size
with probability
, and signed size
with probability
. Then one summand has characteristic function
The sum has characteristic function
If , this tends to
This is still compound Poisson. We have not created a fundamentally new kind of limit. We have only allowed the finitely many surviving jumps to choose between two possible sizes.
The same calculation works for any fixed finite list of jump sizes. Suppose possible magnitudes are , and a signed jump of size
occurs with probability
. Then one summand has characteristic function
The sum has characteristic function
If for each
, then the limit is
This is again compound Poisson, now with finitely many jump classes. A finite list of jump sizes gives a finite sum in the exponent. To get something more like a stable law, we need more and more possible scales.
Now allow the number of possible scales to grow with . Imagine possible magnitudes
where
. Suppose a signed jump of size
occurs with probability
. Then one summand has characteristic function
and the sum has characteristic function
The important quantities are This is the expected number of jumps at scale
among
observations. If only finitely many
remain significant, we fall back to finite compound Poisson behavior. If many tiny scales contribute and individual jumps become negligible, a Gaussian component may appear. If jumps survive across a growing range of scales, we get an infinitely divisible jump law.
Instead of tracking each , ask for the expected number of jumps larger than a threshold
:
For the dyadic model, if , then jumps larger than
are the jumps at levels
. Therefore
So the many-scale problem becomes: how do these cumulative counts grow as we move down the scales?
If the cumulative count stays bounded, only finitely many jumps survive. If it grows very fast but only at tiny scales, those tiny jumps may average into Gaussian behavior. If it grows regularly from scale to scale, then a stable-type hierarchy appears.
Power laws enter as the cleanest scale-regular choice. Suppose that each time we reduce the scale by a factor of , the expected number of visible jumps is multiplied by
. Then the expected number of jumps above scale
grows like
In continuous notation, this becomes
This is the scale-counting reason for power laws. We are not inserting a power law just because stable laws are famous. We are asking for jump counts that transform regularly as we zoom in. Reducing by a factor of
multiplies the expected number of visible jumps by
. To realize this with bounded finite-stage variables, we use a cutoff power-law density. Fix
, and let
. Define a symmetric density on
by
The upper cutoff keeps all variables bounded. The lower cutoff
says that at stage
we include only finitely many scales. As
grows,
decreases, so more small scales enter.
Let us compute the tail. For ,
Substituting the density,
Since we get
Thus for fixed ,
This identifies the central parameter
For rare coins, counted how many jumps of one available size survived. Here
counts how many macroscopic jumps survive in a many-scale distribution.
Moments also reflect the same parameter. For ,
Since we have
In particular, for and
,
Therefore
Thus the same quantity controls both macroscopic jump counts and total variance.
The dense regime: many jumps, Gaussian after standardization
First assume Then for every fixed
,
So there are many jumps above every fixed visible scale. At first this may look very non-Gaussian. But the total variance is also going to infinity:
Let Then
. Since
,
At the standard-deviation scale of the sum, every single jump becomes invisible. The entire interval
collapses to
.
Now compute the characteristic function of . Since the variables are centered and symmetric,
where the remainder can be bounded by a constant times
Multiplying by , the quadratic term gives
For the remainder, since is of order
, while
is of order
, the total error is of order
Therefore
In the dense regime, the many-scale structure exists at the raw level, but the standard-deviation normalization crushes every individual jump. Only the total quadratic variation remains. That is why the Gaussian mechanism returns.
The critical regime: finitely many visible jumps at each fixed scale
Now assume Then for fixed
,
Above each fixed threshold , only finitely many jumps survive. So we should not divide by a scale going to infinity, because doing so would crush the visible jumps. We look at the raw sum. By symmetry,
Multiplying by and using
gives
Hence
This is not Gaussian. A Gaussian exponent is quadratic in . Here the exponent is an integral over jump sizes
. The finite-scale model gave finite sums such as
The many-scale model replaces the finite sum by an integral. The corresponding jump intensity is
This is a Lévy measure. Concretely, it says how many jumps of each size survive. The factor says that smaller jumps are more numerous. The cutoff
remembers the original bounded support. This is why the limit is naturally called a truncated stable law: it has stable-like behavior at small scales, but no jump larger than the original cutoff.
The sparse regime:
Now assume Then for every fixed
,
No jump survives at any fixed macroscopic scale. Also so
in probability. If we look at the original scale, the limit is trivial. But the sum vanished because we were looking too coarsely. The interesting activity has moved below every fixed scale. To find the correct scale, solve for
such that order-one many jumps exceed
:
Using , we get
Thus
Since
we have
. So we magnify the sum and study
. The support changes dramatically. Originally active magnitudes lie in
. After division by
, they lie in
But
So under the stable magnification, the bounded interval of possible magnitudes expands toward . This is how a full unbounded stable law emerges from bounded finite-stage variables.
At the magnified scale, for fixed ,
Thus the number of normalized jumps larger than
has the stable scale-counting form. Now compute the characteristic function. We have
Change variables . Then
, and
So the expression becomes approximately
But , so the prefactor is
. The lower limit tends to
and the upper limit tends to
. Therefore
The integral has the scaling form of a stable law. Indeed, with ,
Thus
So where
is a symmetric
-stable random variable.
The three regimes can now be summarized cleanly. If , many jumps survive at fixed scales, but the standard-deviation scale grows so much that individual jumps disappear; the limit is Gaussian. If
, finitely many jumps survive at fixed scales; the raw sum has a truncated stable jump limit. If
, no fixed-scale jumps survive; the raw sum vanishes, but a smaller scale reveals a full stable law.
What convergence can hide
At finite , every
is bounded. Yet in the sparse stable regime, the normalized sums
may have variances tending to infinity. This is not a contradiction. Convergence in distribution does not automatically imply convergence of moments.
In the sparse stable regime, Therefore
If , then
. Since
, the variance diverges. So
can converge in distribution to a stable random variable while its variances blow up.
A toy example shows the danger. Let with probability
, and
otherwise. Then
in probability, because the exceptional event has probability
. But
If instead
with probability
, then still
in probability, but
Rare events can disappear from ordinary probability convergence while controlling expectations and higher moments. Stable limits are a systematic version of this phenomenon.
For positive stable laws, infinite mean appears similarly. Suppose and use the one-sided density
Each variable satisfies
. In the sparse regime
, with
, the normalized sum
converges to a positive
-stable law. The raw mean satisfies
. But the normalized mean behaves like
Since
, this is
For , the exponent
is negative, so this tends to infinity. Thus the raw sums vanish, while the normalized sums converge to a finite random variable with infinite mean. The infinite mean comes from rare normalized extremes.
Largest Term
Another way to distinguish Gaussian and stable behavior is to look at the maximum
In the dense Gaussian regime, with ,
The largest observation is negligible compared with the scale of the whole sum. This is one of the signatures of Gaussian behavior.
In the sparse stable regime, with , the maximum is on the same scale as the sum. For fixed
,
This equals Since
we get
So in the stable regime, the largest observation remains visible. It is not necessarily the whole sum, because several large jumps may contribute, but it lives on the same scale. This is a major difference between Gaussian and stable addition. In the Gaussian world, the maximum is negligible. In the stable world, extreme terms are part of the main structure.
The Lévy–Khintchine picture
We have now seen several mechanisms: Gaussian behavior appears when many tiny centered contributions accumulate, with no individual term visible. Poisson behavior appears when finitely many visible events survive. Compound Poisson behavior appears when finitely many visible jumps survive, and their sizes are random. Truncated stable behavior appears when jumps survive across infinitely many small scales, but the original upper cutoff remains visible. Full stable behavior appears when the correct normalization magnifies the bounded support so much that the upper cutoff disappears.
The Lévy–Khintchine formula is the general bookkeeping system for all of these mechanisms. Let us build it from examples rather than state it abstractly. Suppose a summand has four independent pieces:
Here is a tiny deterministic displacement,
is a tiny centered fluctuation,
is a rare visible jump with
, and
is a many-scale small-jump variable.
Because the four pieces are independent, the one-step characteristic function approximately factors into four pieces:
Now raise this to the th power and take logarithms. The deterministic part gives
The Gaussian part gives
The rare visible jumps give
The many-scale jumps give an integral over jump sizes. If the limiting jump intensity is a measure , then the jump contribution has the form
with a possible drift correction if the small jumps are asymmetric.
Thus a general limiting exponent has the form
This is the Lévy–Khintchine formula. The characteristic function is
Each term has a concrete meaning.
The term is drift. In the simplest case, if every summand contributes
, then the sum contributes
, and the characteristic function is
. That is pure deterministic motion. In more complicated asymmetric jump models, drift also includes the deterministic first-order part left after small-jump compensation. So drift is not simply “the mean.” It is the net deterministic motion after we have separated visible jumps from their compensated small-jump part. The term
is the Gaussian part. It comes from many individually invisible centered fluctuations whose total variance converges to
. The measure
is the Lévy measure. It records the intensity of jumps of different sizes. If
is a set away from zero, then
can be read as the limiting expected number of jumps with sizes in
.
For compound Poisson jumps with jump distribution and rate
,
Then
which is exactly the compound Poisson exponent. If there are finitely many jump sizes , then
and the jump exponent becomes
If the jumps are symmetric, the imaginary parts cancel and we get
This is exactly the finite-scale calculation.
In the truncated stable regime, the finite sum becomes an integral. The Lévy measure is
Since the measure is symmetric,
Thus
This is the continuum version of the finite jump-size sum. The finite-scale model has a finite list of surviving jump sizes. The stable-type model has a continuum of surviving jump sizes.
There is one important subtlety. In the compound Poisson case, , so the total number of jumps is finite. In the truncated stable case,
So the total mass of near zero is infinite. This means infinitely many tiny jumps accumulate near zero. However, for every fixed
, the mass of
is finite. There are only finitely many jumps larger than any fixed threshold. The infinity is hidden near zero.
This is why the Lévy–Khintchine formula subtracts the linear term . For small
,
If there are infinitely many tiny jumps, the integral of the linear term may not converge. The compensated expression
behaves like near zero, which is much more integrable. The defining condition on a Lévy measure is
This says: there cannot be too many large jumps, and the infinitely many small jumps must be square-summable enough for the characteristic exponent to make sense after compensation.
The drift term is clearest from this compensation viewpoint. In symmetric examples, drift usually vanishes because positive and negative first-order effects cancel. But in asymmetric examples, infinitely many small jumps may have a systematic average direction. The compensation subtracts their linear part from the integral, and the leftover deterministic correction is placed into the drift . Thus drift is the deterministic first-order motion left after centering and small-jump compensation. It is not the same thing as replacing visible jumps by their average. Visible rare jumps remain random and belong in the jump measure.
The Lévy–Khintchine formula is therefore not mysterious. It is a compact record of what can survive after adding many independent pieces:
A pure Gaussian limit occurs when every individual normalized summand is negligible and only quadratic variation remains. A compound Poisson limit occurs when finitely many jumps survive. A truncated stable limit occurs when jumps survive at infinitely many small scales but the upper cutoff remains visible. A full stable limit occurs when normalization magnifies the support so that the cutoff disappears. A drift appears when deterministic first-order motion survives.
A triangular array does not have to choose one clean regime along the whole sequence. The parameter may oscillate. Along subsequences where
, one may see Gaussian behavior. Along subsequences where
, one may see truncated stable behavior. Along subsequences where
, one may see stable behavior after magnification. Limits depend on parameter regimes, not merely on the formula for one
.
There are also borderline corrections. A tail may behave like where
varies slowly with scale. Then the correct normalization may include logarithmic or slowly varying factors. The philosophy remains the same: solve
to find the scale where visible extremes live, then ask whether the maximum is negligible, finite, or dominant, and ask what happens to the small jumps below that scale.
Conclusion
The central limit theorem is not simply the statement that sums of independent variables become Gaussian. More precisely, Gaussian behavior arises when, after normalization, every individual summand becomes negligible and only the aggregate quadratic variation remains. Poisson behavior arises when only finitely many visible events survive. Compound Poisson behavior arises when finitely many visible events survive and carry random sizes. Truncated stable behavior arises when jumps survive at a continuum of small scales, but the original upper cutoff remains visible. Full stable behavior arises when the raw sum vanishes at fixed scale, but after magnifying to the scale where order-one jumps reappear, the bounded support expands to an unbounded range.
The power-law construction was chosen because we wanted jump counts that behave regularly across scales:
This leads to densities proportional to , with a lower cutoff
and an upper cutoff
. The phase diagram is governed by
If
, Gaussian behavior returns after standardization. If
, the raw sum has a truncated stable jump limit. If
, the raw sum vanishes, but the magnified sum has a full stable limit.
The final moral is this: boundedness alone does not decide the limiting behavior of sums. What matters is what remains visible at the scale where we observe the sum. The normalization decides which events are crushed to zero, which jumps remain visible, which rare tail events are magnified, and whether the limit records drift, Gaussian fluctuation, jumps, or some combination of all three.
.