Poles and Zeros: What Comes Out Uninvited, and What Goes In and Vanishes

Why exponentials are the natural test input for a linear system, the physical reading of a pole and a zero, and how a plot of two sets of points summarizes a whole system.

Study / AI Theory / Signals, Systems & Transforms / Representations & System Response
X LinkedIn

A pole–zero plot compresses a system’s algebra, but it does not by itself specify causality, stability, or the response to initial conditions. Those conclusions require the region of convergence and a clear statement of the system being represented.

The organizing calculation is to feed an LTI system an exponential. The z-transform appears as its multiplier. Difference equations then explain why poles describe natural modes, while convolution explains why zeros can suppress particular inputs.

From an exponential input to the z-transform

For a discrete-time LTI system with impulse response \(h[n]\),

\[y[n]=\sum_{k=-\infty}^{\infty}h[k]x[n-k].\]

Choose an all-time exponential,

\[x[n]=\alpha^n, \qquad \alpha\ne0.\]

Substitution gives

\[\begin{aligned} y[n] &= \sum_k h[k]\alpha^{n-k}\\ &= \alpha^n\sum_k h[k]\alpha^{-k}. \end{aligned}\]

Define

\[H(z)=\sum_{k=-\infty}^{\infty}h[k]z^{-k}.\]

Then, whenever the sum converges,

\[y[n]=H(\alpha)\alpha^n.\]

The negative exponent in the transform is not arbitrary decoration. A delay contributes

\[\alpha^{n-k}=\alpha^n\alpha^{-k},\]

so the transform collects exactly the factors generated by delayed copies of the input.

This is a bilateral transform: indices range through both positive and negative time. The associated region of convergence, abbreviated ROC, is part of the transform pair.

A switched-on exponential,

\[x[n]=\alpha^n u[n],\]

is a different input. Near startup, the convolution lacks contributions from negative-time input samples. Its response generally contains additional terms, even when the corresponding all-time exponential is an eigenfunction.

Why delays turn difference equations into algebra

Let

\[X(z)=\sum_n x[n]z^{-n}.\]

For a delayed sequence, substitute a new summation index:

\[\begin{aligned} \mathcal Z\{x[n-r]\} &= \sum_n x[n-r]z^{-n}\\ &= \sum_m x[m]z^{-(m+r)}\\ &= z^{-r}X(z). \end{aligned}\]

Now consider a constant-coefficient difference equation,

\[y[n]+\sum_{r=1}^{P}a_r y[n-r] = \sum_{r=0}^{Q}b_r x[n-r].\]

For the zero-state transfer relation,

\[\left(1+\sum_{r=1}^{P}a_r z^{-r}\right)Y(z) = \left(\sum_{r=0}^{Q}b_r z^{-r}\right)X(z),\]

so

\[H(z)=\frac{Y(z)}{X(z)} = \frac{\sum_{r=0}^{Q}b_r z^{-r}} {1+\sum_{r=1}^{P}a_r z^{-r}}.\]

Finite-order constant-coefficient equations produce rational transfer functions. An arbitrary LTI system need not have a rational transform, so “the transfer function is a ratio of finite polynomials” is a statement about this important system class, not every possible LTI operator.

For a rational function in reduced form, zeros are roots of the numerator and poles are roots of the denominator. Common factors must be examined before classifying them.

A first-order system calculated three ways

Take

\[y[n]-\frac12 y[n-1]=x[n].\]

Assume initial rest. For an impulse input, the recurrence gives

\[h[0]=1, \qquad h[n]=\frac12h[n-1]\quad(n\ge1).\]

Therefore,

\[h[n]=\left(\frac12\right)^n u[n].\]

Its transform is a geometric series:

\[\begin{aligned} H(z) &= \sum_{n=0}^{\infty} \left(\frac12z^{-1}\right)^n\\ &= \frac{1}{1-\frac12z^{-1}}, \qquad \lvert z\rvert>\frac12. \end{aligned}\]

Multiplying numerator and denominator by \(z\) gives

\[H(z)=\frac{z}{z-\frac12}.\]

The pole lies at one half. The expression also has a zero at the origin, but zero is outside this ROC and does not specify an admissible all-time exponential input. Algebraic roots and physically meaningful test inputs should not be conflated.

For a unit step input, convolution gives

\[\begin{aligned} y[n] &= \sum_{k=0}^{n}\left(\frac12\right)^k\\ &= \frac{1-(\frac12)^{n+1}}{1-\frac12}\\ &= 2-\left(\frac12\right)^n, \qquad n\ge0. \end{aligned}\]

The first outputs are

\[1,\quad \frac32,\quad \frac74,\quad \frac{15}{8}.\]

They approach the DC gain,

\[H(1)=2.\]

The complete step response still differs from the constant output at every finite time. The pole describes that difference:

\[y[n]-2=-\left(\frac12\right)^n.\]

For a nonzero initial condition,

\[y[-1]=c,\]

the recurrence instead yields

\[y[n]=2+(c-2)\left(\frac12\right)^{n+1}.\]

Checking the first output,

\[y[0]=2+\frac{c-2}{2}=1+\frac c2,\]

which agrees directly with the difference equation.

Poles and natural modes

Set the input to zero in a general difference equation and try

\[y[n]=p^n.\]

Substitution gives

\[p^n+\sum_{r=1}^{P}a_r p^{n-r}=0.\]

For nonzero \(p\), division by \(p^{n-P}\) produces the characteristic polynomial,

\[p^P+a_1p^{P-1}+\cdots+a_P=0.\]

These roots are the candidate natural modes. In a minimal causal realization, the uncancelled poles of the transfer function correspond to these modes.

For a simple pole,

\[p=\rho e^{j\theta},\]

the associated exponential is

\[p^n=\rho^n e^{j\theta n}.\]

Its magnitude decays when the radius is below one, remains constant when the radius is one, and grows when the radius exceeds one.

Repeated poles require more care. Consider

\[H(z)=\frac{1}{(1-a z^{-1})^2}\]

with the causal ROC. This is the product of two identical first-order transforms, so its impulse response is the convolution

\[\begin{aligned} h[n] &= \sum_{k=0}^{n}a^k a^{n-k}\\ &= (n+1)a^n. \end{aligned}\]

The repeated pole creates a polynomial factor. At unit radius, that factor can grow even though a single exponential would have constant magnitude.

This is one reason the causal BIBO stability condition requires poles strictly inside the unit circle. “Not outside” is insufficient.

The same rational expression can describe different sequences

Consider

\[X(z)=\frac{1}{1-a z^{-1}}, \qquad a\ne0.\]

For the right-sided sequence,

\[x_{\mathrm R}[n]=a^n u[n],\]

the transform is

\[\sum_{n=0}^{\infty}(a/z)^n,\]

which converges when

\[\lvert z\rvert>\lvert a\rvert.\]

Now take a left-sided sequence,

\[x_{\mathrm L}[n]=-a^n u[-n-1].\]

Its transform is

\[\begin{aligned} X_{\mathrm L}(z) &= -\sum_{n=-\infty}^{-1}a^n z^{-n}\\ &= -\sum_{m=1}^{\infty}(z/a)^m\\ &= -\frac{z/a}{1-z/a}\\ &= \frac{1}{1-a z^{-1}}, \end{aligned}\]

but convergence now requires

\[\lvert z\rvert<\lvert a\rvert.\]

The algebraic expression is identical; the sequences are not.

For example, set

\[a=2.\]

The causal sequence grows toward the future and is not absolutely summable. The left-sided sequence decays toward the past, and

\[\sum_{n=-\infty}^{-1}\lvert -2^n\rvert = \sum_{m=1}^{\infty}2^{-m} = 1.\]

Thus a stable noncausal system can have a pole outside the unit circle. The rule “all poles must be inside” requires causality and a rational transfer function.

Why the ROC is an annulus

For the bilateral transform, split the sum into positive and negative indices. The positive-time part places a lower bound on the allowed radius. The negative-time part places an upper bound.

Their intersection has the form

\[r_{\mathrm{in}}<\lvert z\rvert<r_{\mathrm{out}},\]

with possible inclusion of appropriate boundaries, the origin, or infinity depending on the sequence. The open annular interior contains no poles.

For a rational right-sided sequence, the ROC extends outward beyond the outermost pole. Causality additionally requires no negative-time samples; for rational transforms this is reflected in the behavior at infinity. A finite advance, for example, should not be mistaken for a causal sequence merely because it has no ordinary finite poles.

A useful two-sided example is

\[h[n] = \left(\frac12\right)^n u[n] - 2^n u[-n-1].\]

Its transform is

\[H(z) = \frac{1}{1-\frac12z^{-1}} + \frac{1}{1-2z^{-1}},\]

with ROC

\[\frac12<\lvert z\rvert<2.\]

The unit circle lies inside this annulus. Directly,

\[\sum_n\lvert h[n]\rvert = \sum_{n=0}^{\infty}\left(\frac12\right)^n + \sum_{m=1}^{\infty}2^{-m} = 2+1=3.\]

This example is stable and noncausal, with poles on both sides of the unit circle.

Deriving the BIBO stability condition

Suppose the input is bounded:

\[\lvert x[n]\rvert\le B.\]

Then

\[\begin{aligned} \lvert y[n]\rvert &= \left|\sum_k h[k]x[n-k]\right|\\ &\le \sum_k\lvert h[k]\rvert\lvert x[n-k]\rvert\\ &\le B\sum_k\lvert h[k]\rvert. \end{aligned}\]

Absolute summability of the impulse response therefore guarantees a finite output bound.

The converse can be understood by choosing input phases to align with the impulse response. At one output time, select bounded input values so each product contributes positively:

\[x[-k]= \begin{cases} \overline{h[k]}/\lvert h[k]\rvert,&h[k]\ne0,\\ 0,&h[k]=0. \end{cases}\]

Finite truncations of this input produce output values equal to partial sums of

\[\sum_k\lvert h[k]\rvert.\]

If those sums grow without bound, no common bounded-input output bound exists.

For ordinary discrete-time convolution systems,

\[\text{BIBO stability} \iff \sum_k\lvert h[k]\rvert<\infty.\]

Under the absolute-convergence definition of the ROC, this is equivalent to the unit circle being included. A Fourier transform can sometimes exist conditionally or in a generalized sense without absolute summability; that weaker existence does not establish BIBO stability.

Zeros as annihilated exponential inputs

If a nonzero complex number belongs to the ROC and satisfies

\[H(\alpha)=0,\]

then the admissible all-time input

\[x[n]=\alpha^n\]

has zero output.

For a real system, a unit-circle zero generally comes with its conjugate. Together they suppress the corresponding real sinusoid.

Consider the four-tap moving sum:

\[y[n]=x[n]+x[n-1]+x[n-2]+x[n-3].\]

Its response is

\[H(z)=1+z^{-1}+z^{-2}+z^{-3}.\]

Using a geometric-series identity,

\[H(z)=\frac{1-z^{-4}}{1-z^{-1}}.\]

The apparent singularity at one is removable:

\[H(1)=1+1+1+1=4.\]

The genuine zeros are

\[z=-1,\qquad z=j,\qquad z=-j.\]

There are three, not four. The fourth root of unity at one was cancelled.

For the all-time periodic sequence

\[x[n]=\cos\left(\frac{\pi n}{2}\right),\]

every four-sample window contains a cyclic permutation of

\[1,\quad0,\quad-1,\quad0.\]

Its sum is exactly zero.

Switching the same sinusoid on at zero changes the first outputs. From rest they are

\[1,\quad1,\quad0,\quad0,\quad0,\ldots\]

because the first windows contain missing past samples. A zero suppresses the complete exponential mode; it does not erase every startup effect associated with switching that mode on.

The analogous ten-tap moving sum has nine unit-circle zeros: the tenth roots of unity other than one.

Reading frequency response from distances

Write a reduced rational function as

\[H(z)=C \frac{\prod_{r=1}^{R}(z-z_r)} {\prod_{q=1}^{Q}(z-p_q)}.\]

At a point on the unit circle,

\[\lvert H(e^{j\omega})\rvert = \lvert C\rvert \frac{\prod_r\lvert e^{j\omega}-z_r\rvert} {\prod_q\lvert e^{j\omega}-p_q\rvert}.\]

Each numerator factor is a distance to a zero; each denominator factor is a distance to a pole.

The phase is the corresponding sum of angles:

\[\arg H(e^{j\omega}) = \arg C + \sum_r\arg(e^{j\omega}-z_r) - \sum_q\arg(e^{j\omega}-p_q).\]

This derives the geometric reading of a pole–zero plot. Nearby zeros reduce the numerator, and nearby poles reduce the denominator. Other factors and cancellations still matter, so proximity alone is not a complete specification.

For example,

\[H(z)=\frac{1+z^{-1}}{1-\frac12z^{-1}} = \frac{z+1}{z-\frac12}.\]

It has a zero at negative one and a pole at one half. For the causal ROC,

\[H(1)=4, \qquad H(-1)=0.\]

At a quarter of the sampling frequency,

\[\begin{aligned} H(j) &= \frac{1-j}{1+\frac12j}\\ &= \frac{(1-j)(1-\frac12j)}{1+\frac14}\\ &= \frac25-\frac65j. \end{aligned}\]

Therefore,

\[\lvert H(j)\rvert = \sqrt{\frac85}.\]

The same result follows from distances:

\[\frac{\lvert j+1\rvert}{\lvert j-\frac12\rvert} = \frac{\sqrt2}{\sqrt{5/4}} = \sqrt{\frac85}.\]

This is a direct check that the algebraic and geometric readings agree.

Pole cancellation and initial conditions

Consider the equation

\[y[n]-2y[n-1]=x[n]-2x[n-1].\]

The zero-state transfer function appears to be

\[H(z)=\frac{1-2z^{-1}}{1-2z^{-1}}=1.\]

But define the mismatch

\[e[n]=y[n]-x[n].\]

The original recurrence gives

\[e[n]=2e[n-1].\]

An initial mismatch grows even though the reduced zero-state input–output transfer function is unity. Exact cancellation can hide a mode of a nonminimal realization.

Perturb the numerator slightly:

\[H_\varepsilon(z) = \frac{1-(2-\varepsilon)z^{-1}}{1-2z^{-1}} = 1+\frac{\varepsilon z^{-1}}{1-2z^{-1}}.\]

The causal impulse response becomes

\[h_\varepsilon[n] = \delta[n]+\varepsilon 2^{n-1}u[n-1].\]

Any nonzero perturbation exposes the unstable tail. This is an algebraic construction, not a numerical sensitivity measurement.

Initial conditions can also be retained directly through a unilateral transform. For

\[Y_+(z)=\sum_{n=0}^{\infty}y[n]z^{-n},\]

the delayed sequence transforms as

\[\sum_{n=0}^{\infty}y[n-1]z^{-n} = y[-1]+z^{-1}Y_+(z).\]

Thus a first-order recurrence retains its initial value explicitly. The transfer function does not eliminate initial conditions; a zero-state interpretation sets them aside.

Revision checklist

Check What I should be able to reproduce
Transform definition Derive the negative powers from delayed exponential inputs.
Difference equation Convert each delay into a power of the transform variable.
First-order response Calculate both the impulse response and step response.
Natural modes Derive the characteristic polynomial from a homogeneous exponential.
Repeated poles Obtain the polynomial factor by convolving two geometric sequences.
ROC ambiguity Derive both inverses of the same first-order rational expression.
Stability Prove the absolute-summability bound and explain its converse.
Unit-circle zeros Sum one complete period and account for startup separately.
Plot geometry Derive magnitude as a ratio of products of distances.
Cancellation Explain why reduced zero-state behavior can hide an unstable mode.

Why it matters for my work

Poles and zeros provide exact questions for linear preprocessing and local dynamical models: which disturbances decay, which frequencies disappear, and which assumptions make those conclusions valid? For nonlinear learned systems, the vocabulary can guide experiments, but the linear guarantees require a justified linear model.

What I have not resolved

For a deployed signal pipeline, I still need to distinguish the mathematical filter from its finite-precision realization and startup policy. A stable reduced transfer function is not enough to establish that every implementation behaves safely under initialization and coefficient perturbations.

Related study notes

← Back to AI Theory