System size expansion explained

The system size expansion, also known as van Kampen's expansion or the Ω-expansion, is a technique pioneered by Nico van Kampen^[1] used in the analysis of stochastic processes. Specifically, it allows one to find an approximation to the solution of a master equation with nonlinear transition rates. The leading order term of the expansion is given by the linear noise approximation, in which the master equation is approximated by a Fokker–Planck equation with linear coefficients determined by the transition rates and stoichiometry of the system.

Less formally, it is normally straightforward to write down a mathematical description of a system where processes happen randomly (for example, radioactive atoms randomly decay in a physical system, or genes that are expressed stochastically in a cell). However, these mathematical descriptions are often too difficult to solve for the study of the systems statistics (for example, the mean and variance of the number of atoms or proteins as a function of time). The system size expansion allows one to obtain an approximate statistical description that can be solved much more easily than the master equation.

Preliminaries

P(X,t)

, giving the probability of observing the system in state

at time

may be, for example, a vector with elements corresponding to the number of molecules of different chemical species in a system. In a system of size

\Omega

(intuitively interpreted as the volume), we will adopt the following nomenclature:

is a vector of macroscopic copy numbers,

x=X/\Omega

is a vector of concentrations, and

\phi

is a vector of deterministic concentrations, as they would appear according to the rate equation in an infinite system.

and

are thus quantities subject to stochastic effects.

A master equation describes the time evolution of this probability. Henceforth, a system of chemical reactions^[2] will be discussed to provide a concrete example, although the nomenclature of "species" and "reactions" is generalisable. A system involving

species and

reactions can be described with the master equation:

	\partialP(X,t)
	\partialt

=\Omega

	R
\sum
	j=1

\left(

	N
\prod
	i=1

	-S_ij
E

-1\right)f_j(x,\Omega)P(X,t).

Here,

\Omega

is the system size,

is an operator which will be addressed later,

S_ij

is the stoichiometric matrix for the system (in which element

S_ij

gives the stoichiometric coefficient for species

in reaction

), and

f_j

is the rate of reaction

given a state

and system size

\Omega

	-S_ij
E

is a step operator, removing

S_ij

from the

th element of its argument. For example,

	-S₂₃
E

f(x_1,x_2,x₃₎=f(x_1,x₂-S₂₃,x₃₎

. This formalism will be useful later.

The above equation can be interpreted as follows. The initial sum on the RHS is over all reactions. For each reaction

, the brackets immediately following the sum give two terms. The term with the simple coefficient −1 gives the probability flux away from a given state

due to reaction

changing the state. The term preceded by the product of step operators gives the probability flux due to reaction

changing a different state

into state

. The product of step operators constructs this state

Example

For example, consider the (linear) chemical system involving two chemical species

X₁

and

X₂

and the reaction

X₁ → X₂

. In this system,

N=2

(species),

R=1

(reactions). A state of the system is a vector

X=\{n_1,n₂\}

, where

n_1,n₂

are the number of molecules of

X₁

and

X₂

respectively. Let

f_1(x,\Omega)=

	n₁
	\Omega

=x₁

, so that the rate of reaction 1 (the only reaction) depends on the concentration of

X₁

. The stoichiometry matrix is

(-1,1)^T

Then the master equation reads:

\begin{align}

	\partialP(X,t)
	\partialt

&=\Omega\left(

	-S₁₁
E

	-S₂₁
E

-1\right)f₁\left(

	X
	\Omega

\right)P(X,t)\\ &=\Omega\left(f₁\left(

	X+\DeltaX
	\Omega

\right)P\left(X+\DeltaX,t\right)-f₁\left(

	X
	\Omega

\right)P\left(X,t\right)\right),\end{align}

where

\DeltaX=\{1,-1\}

is the shift caused by the action of the product of step operators, required to change state

to a precursor state

Linear noise approximation

If the master equation possesses nonlinear transition rates, it may be impossible to solve it analytically. The system size expansion utilises the ansatz that the variance of the steady-state probability distribution of constituent numbers in a population scales like the system size. This ansatz is used to expand the master equation in terms of a small parameter given by the inverse system size.

Specifically, let us write the

X_i

, the copy number of component

, as a sum of its "deterministic" value (a scaled-up concentration) and a random variable

\xi

, scaled by

\Omega^1/2

X_i=\Omega\phi_i+\Omega^1/2\xi_i.

The probability distribution of

can then be rewritten in the vector of random variables

\xi

P(X,t)=P(\Omega\phi+\Omega^1/2\xi)=\Pi(\xi,t).

Consider how to write reaction rates

and the step operator

in terms of this new random variable. Taylor expansion of the transition rates gives:

f_j(x)=f_j(\phi+\Omega^-1/2\xi)=f_j(\phi)+\Omega^-1/2

	N
\sum
	i=1

	\partialf_j(\phi)
	\partial\phi_i

\xi_i+O(\Omega^-1).

The step operator has the effect

Ef(n) → f(n+1)

and hence

Ef(\xi) → f(\xi+\Omega^-1/2)

	N
\prod
	i=1

	-S_ij
E

\simeq1-\Omega^-1/2\sum_iS_ij

	\partial
	\partial\xi_i

	\Omega^-1
	2

\sum_i\sum_kS_ijS_kj

	\partial²
	\partial\xi_i\partial\xi_k

+O(\Omega^-3/2).

We are now in a position to recast the master equation.

\begin{align}&{}

	\partial\Pi(\xi,t)
	\partialt

-\Omega^1/2

	N
\sum
	i=1

	\partial\phi_i
	\partialt

	\partial\Pi(\xi,t)
	\partial\xi_i

\\ &=\Omega

	R
\sum
	j=1

\left(-\Omega^-1/2\sum_iS_ij

	\partial
	\partial\xi_i

	\Omega^-1
	2

\sum_i\sum_kS_ijS_kj

	\partial²
	\partial\xi_i\partial\xi_k

+O(\Omega^-3/2)\right)\\ &{} x \left(f_j(\phi)+\Omega^-1/2\sum_i

	\partialf_j(\phi)
	\partial\phi_i

\xi_i+O(\Omega^-1)\right)\Pi(\xi,t).\end{align}

This rather frightening expression makes a bit more sense when we gather terms in different powers of

\Omega

. First, terms of order

\Omega^1/2

give

	N
\sum
	i=1

	\partial\phi_i
	\partialt

	\partial\Pi(\xi,t)
	\partial\xi_i

	N
\sum
	i=1

	R
\sum
	j=1

S_ijf_j(\phi)

	\partial\Pi(\xi,t)
	\partial\xi_i

These terms cancel, due to the macroscopic reaction equation

	\partial\phi_i
	\partialt

	R
\sum
	j=1

S_ijf_j(\phi).

The terms of order

\Omega⁰

are more interesting:

	\partial\Pi(\xi,t)
	\partialt

=\sum_j\left(\sum_ik-S_ij

	\partialf_j
	\partial\phi_k

	\partial(\xi_k\Pi(\xi,t))
	\partial\xi_i

	1
	2

f_j\sum_ikS_ijS_kj

	\partial²\Pi(\xi,t)
	\partial\xi_i\partial\xi_k

\right),

which can be written as

	\partial\Pi(\xi,t)
	\partialt

=-\sum_ikA_ik

	\partial(\xi_k\Pi)
	\partial\xi_i

	1
	2

\sum_ik

	T]
[BB
	ik

	\partial²\Pi
	\partial\xi_i\partial\xi_k

where

A_ik=

	R
\sum
	j=1

S_ij

	\partialf_j
	\partial\phi_k

	\partial(S_i ⋅ f)
	\partial\phi_k

and

[BB^T]_ik=

	R
\sum
	j=1

S_ijS_kjf_j(\phi)=[Sdiag(f(\phi))S^T]_ik.

The time evolution of

\Pi

is then governed by the linear Fokker–Planck equation with coefficient matrices

and

BB^T

(in the large-

\Omega

limit, terms of

O(\Omega^-1/2)

may be neglected, termed the linear noise approximation). With knowledge of the reaction rates

and stoichiometry

, the moments of

\Pi

can then be calculated.

The approximation implies that fluctuations around the mean are Gaussian distributed. Non-Gaussian features of the distributions can be computed by taking into account higher order terms in the expansion.^[3]

Software

The linear noise approximation has become a popular technique for estimating the size of intrinsic noise in terms of coefficients of variation and Fano factors for molecular species in intracellular pathways. The second moment obtained from the linear noise approximation (on which the noise measures are based) are exact only if the pathway is composed of first-order reactions. However bimolecular reactions such as enzyme-substrate, protein-protein and protein-DNA interactions are ubiquitous elements of all known pathways; for such cases, the linear noise approximation can give estimates which are accurate in the limit of large reaction volumes. Since this limit is taken at constant concentrations, it follows that the linear noise approximation gives accurate results in the limit of large molecule numbers and becomes less reliable for pathways characterized by many species with low copy numbers of molecules.

The system size expansion and linear noise approximation have been made available via automated derivation in an open source software project Multi-Scale Modelling Tool (MuMoT).^[4]

A number of studies have elucidated cases of the insufficiency of the linear noise approximation in biological contexts by comparison of its predictions with those of stochastic simulations.^[5] ^[6] This has led to the investigation of higher order terms of the system size expansion that go beyond the linear approximation. These terms have been used to obtain more accurate moment estimates for the mean concentrations and for the variances of the concentration fluctuations in intracellular pathways. In particular, the leading order corrections to the linear noise approximation yield corrections of the conventional rate equations.^[7] Terms of higher order have also been used to obtain corrections to the variances and covariances estimates of the linear noise approximation.^[8] ^[9] The linear noise approximation and corrections to it can be computed using the open source software intrinsic Noise Analyzer. The corrections have been shown to be particularly considerable for allosteric and non-allosteric enzyme-mediated reactions in intracellular compartments.

Notes and References

van Kampen, N. G. (2007) "Stochastic Processes in Physics and Chemistry", North-Holland Personal Library
Elf, J. and Ehrenberg, M. (2003) "Fast Evaluation of Fluctuations in Biochemical Networks With the Linear Noise Approximation", Genome Research, 13:2475–2484.
Thomas. Philipp. Grima. Ramon. 2015-07-13. Approximate probability distributions of the master equation. Physical Review E. 92. 1. 012120. 10.1103/PhysRevE.92.012120. 26274137. 1411.3551. 2015PhRvE..92a2120T. 13700533.
Marshall . James A. R. . Reina . Andreagiovanni . Bose . Thomas . Multiscale Modelling Tool: Mathematical modelling of collective behaviour without the maths . PLOS ONE . 30 September 2019 . 14 . 9 . e0222906 . 10.1371/journal.pone.0222906. free . 31568526 . 6768458 . 2019PLoSO..1422906M .
Hayot, F. and Jayaprakash, C. (2004), "The linear noise approximation for molecular fluctuations within cells", Physical Biology, 1:205
Ferm, L. Lötstedt, P. and Hellander, A. (2008), "A Hierarchy of Approximations of the Master Equation Scaled by a Size Parameter", Journal of Scientific Computing, 34:127
Grima, R. (2010) "An effective rate equation approach to reaction kinetics in small volumes: Theory and application to biochemical reactions in nonequilibrium steady-state conditions", The Journal of Chemical Physics, 132:035101
Grima, R. and Thomas, P. and Straube, A.V. (2011), "How accurate are the nonlinear chemical Fokker-Planck and chemical Langevin equations?", The Journal of Chemical Physics, 135:084103
Grima, R. (2012), "A study of the accuracy of moment-closure approximations for stochastic chemical kinetics", The Journal of Chemical Physics, 136: 154105