Jump to content

Method of moments (statistics)

From Wikipedia, the free encyclopedia

In statistics, the method of moments is a method of estimation of population parameters. The same principle is used to derive higher moments like skewness and kurtosis.

It starts by expressing the population moments (i.e., the expected values of powers of the random variable under consideration) as functions of the parameters of interest. Those expressions are then set equal to the sample moments. The number of such equations is the same as the number of parameters to be estimated. Those equations are then solved for the parameters of interest.[1] The solutions are estimates of those parameters. The method of moments was first introduced by Karl Pearson in 1895.[2][3]

Method

[edit]

Suppose that the parameter = () characterizes the distribution of the random variable .[4] Suppose the first moments of the true distribution (the "population moments") can be expressed as functions of the s:

Suppose a sample of size is drawn, resulting in the values . For , let be the j-th sample moment, an estimate of . The method of moments estimator for denoted by is defined to be a solution (if one exists) to the equations:[3]

The method described here for single random variables generalizes in an obvious manner to multiple random variables leading to multiple choices for moments to be used. Different choices generally lead to different solutions.[5][6]

Properties

[edit]

If it exists, a moment estimator based on an independent sample of observations drawn from the distribution has the following statistical properties.

Consistency

[edit]

Provided that the function is invertible and continuous, and the first moments exist (), the moment estimator is strongly consistent. This is a direct consequence of the strong law of large numbers.[7]

In the typical case, where the second moment does not explode as goes to infinity, the above implies (by the Vitali convergence theorem) that the moment estimator is also unbiased in the limit, that is,

Asymptotic normality

[edit]

If the function is even differentiable and , then the moment estimator is also asymptotically normal in the sense[7]where is the covariance matrix of . This is a consequence of the central limit theorem and the delta method.

Advantages and disadvantages

[edit]

On the one hand, the method of moments is fairly simple and broadly applicable, since it does not require any assumptions on the data distribution besides existence of the first few moments. This is particularly useful in settings, where the general shape of the distribution may not be known (e.g. utility functions). Moreover, the resulting estimators are almost always consistent and asymptotically unbiased.[7]

On the other hand, the moment equations do not always have a solution, in which case the method does not provide any estimator. In other cases, there might exist multiple solutions. In view of parameter estimation in a parametric model, moment estimators are, in general, not asymptotically efficient, in contrast to maximum likelihood estimators.[8] Moreover, the method does not necessarily lead to sufficient estimators, i.e., they sometimes fail to take into account all relevant information in the sample.[9]

Alternative method of moments

[edit]

The equations to be solved in the method of moments (MoM) are in general nonlinear and there are no generally applicable guarantees that tractable solutions exist[citation needed]. But there is an alternative approach to using sample moments to estimate data model parameters in terms of known dependence of model moments on these parameters, and this alternative requires the solution of only linear equations or, more generally, tensor equations. This alternative is referred to as the Bayesian-Like MoM (BL-MoM), and it differs from the classical MoM in that it uses optimally weighted sample moments. Considering that the MoM is typically motivated by a lack of sufficient knowledge about the data model to determine likelihood functions and associated a posteriori probabilities of unknown or random parameters, it is odd that there exists a type of MoM that is Bayesian-Like. But the particular meaning of Bayesian-Like leads to a problem formulation in which required knowledge of a posteriori probabilities is replaced with required knowledge of only the dependence of model moments on unknown model parameters, which is exactly the knowledge required by the traditional MoM.[5][6][10][11] The BL-MoM also uses knowledge of a priori probabilities of the parameters to be estimated, when available, but otherwise uses uniform priors.[citation needed]

The BL-MoM has been reported on in only the applied statistics literature in connection with parameter estimation and hypothesis testing using observations of stochastic processes for problems in Information and Communications Theory and, in particular, communications receiver design in the absence of knowledge of likelihood functions or associated a posteriori probabilities[12] and references therein. In addition, the restatement of this receiver design approach for stochastic process models as an alternative to the classical MoM for any type of multivariate data is available in tutorial form at the university website.[13] The applications in[12] and references demonstrate some important characteristics of this alternative to the classical MoM, and a detailed list of relative advantages and disadvantages is given in,[13] but the literature is missing direct comparisons in specific applications of the classical MoM and the BL-MoM.[citation needed]

Examples

[edit]

Poisson distribution

[edit]

Let be Poisson distributed with parameter . The first moment (that is, the mean) is

Since this is already an explicit equation for , the moment estimator is simply the sample mean.

Normal distribution

[edit]

Let be normally distributed with unknown parameters and . Since two parameters must be determined, two equations are needed. The first two moments are given by

Solving the equations for the parameters yields

Given a sample , the corresponding estimators are

which are the sample mean and the (biased) sample variance. In this case, they coincide with the maximum likelihood estimators.

Uniform distribution

[edit]

Consider the uniform distribution on the interval , , with unknown parameters . If then we have[9]

Solving these equations gives

Given a set of samples we can use the sample moments and in these formulae in order to estimate and .

Note, however, that this method can produce inconsistent results in some cases. For example, the set of samples results in the estimate , . Since it is impossible for the set to have been drawn from in this case.

Binomial distribution

[edit]

Let be binomially distributed with unknown parameters and . The first two moments read as[14]

Solving these equations for the parameters yields

Inserting the estimators and for and gives the corresponding moment estimators and .

This is a negative example, where the moment estimators have multiple issues: First, the estimator does not necessarily provide integers, although this can be easily resolved by rounding to the nearest integer.[14] A more serious issue is that both estimators can lead to negative estimates, namely when , which is equivalent to the sample variance being larger than the sample mean . This can happen, for example, in small data sets with large variation. In such cases, the moment estimators are useless.

Nonparametric estimation of mean and variance

[edit]

The method of moment does not actually need any assumptions on the data distribution besides the existence of a certain number of moments. For the estimation of the mean and the variance using observations , assuming only , the very same derivation as for the normal distribution (see above) applies, yielding the estimators

Estimation of polynomial probability densities

[edit]

Another example application of the method of moments is to estimate polynomial probability density distributions. In this case, an approximating polynomial of order is defined on an interval . The method of moments then yields a system of equations, whose solution involves the inversion of a Hankel matrix.[15]

See also

[edit]
[edit]

References

[edit]
  1. Dodge, Yadolah (2008). The Concise Encyclopedia of Statistics. Springer. pp. 348–349. ISBN 978-0-387-32833-1.
  2. Pearson, Karl (1895). "Contributions to the Mathematical Theory of Evolution. II. Skew Variation in Homogeneous Material". Philosophical Transactions of the Royal Society of London A. 186: 343–414.
  3. 1 2 Pearson, Karl (June 1936). "Method of Moments and Method of Maximum Likelihood". Biometrika. 28 (1/2): 34. doi:10.2307/2334123.
  4. Bowman, Kimiko O.; L. R., Shenton (1998). "Estimator: Method of Moments". Encyclopedia of Statistical Sciences. Wiley. pp. 2092–2098.
  5. 1 2 Quandt, Richard E.; Ramsey, James B. (December 1978). "Estimating Mixtures of Normal Distributions and Switching Regressions". Journal of the American Statistical Association. 73 (364): 730. doi:10.2307/2286266.
  6. 1 2 Lindsay, Bruce G.; Basak, Prasanta (June 1993). "Multivariate Normal Mixtures: A Fast Consistent Method of Moments". Journal of the American Statistical Association. 88 (422): 468. doi:10.2307/2290326.
  7. 1 2 3 Shao, Jun (2003). Mathematical Statistics (2nd ed.). Springer. p. 207. ISBN 978-0-387-95382-3.
  8. Dodge, Yadolah (2008). The Concise Encyclopedia of Statistics. Springer. p. 349. ISBN 978-0-387-32833-1.
  9. 1 2 Shao, Jun (2003). Mathematical Statistics (2nd ed.). Springer. p. 208. ISBN 978-0-387-95382-3.
  10. Hansen, Lars Peter (July 1982). "Large Sample Properties of Generalized Method of Moments Estimators". Econometrica. 50 (4): 1029. doi:10.2307/1912775.
  11. Lindsay, Bruce (December 1982). "Conditional Score Functions: Some Optimality Results". Biometrika. 69 (3): 503. doi:10.2307/2335985.
  12. 1 2 Gardner, William (May 1981). "Design of nearest prototype signal classifiers (Corresp.)". IEEE Transactions on Information Theory. 27 (3): 368–372. doi:10.1109/TIT.1981.1056334. ISSN 0018-9448.
  13. 1 2 Cyclostationarity, page 11.4
  14. 1 2 Shao, Jun (2003). Mathematical Statistics (2nd ed.). Springer. p. 209. ISBN 978-0-387-95382-3.
  15. J. Munkhammar, L. Mattsson, J. Rydén (2017) "Polynomial probability distribution estimation using the method of moments". PLoS ONE 12(4): e0174573. https://doi.org/10.1371/journal.pone.0174573