Computational statistics

Computational statistics, or statistical computing, is the bond between

statistical education.^[1]

As in

statistical methods, such as cases with very large sample size and non-homogeneous data sets.^[2]

The terms 'computational statistics' and 'statistical computing' are often used interchangeably, although Carlo Lauro (a former president of the International Association for Statistical Computing) proposed making a distinction, defining 'statistical computing' as "the application of computer science to statistics", and 'computational statistics' as "aiming at the design of algorithm for implementing statistical methods on computers, including the ones unthinkable before the computer age (e.g.

simulation), as well as to cope with analytically intractable problems" [sic].^[3]

The term 'Computational statistics' may also be used to refer to computationally intensive statistical methods including

artificial neural networks and generalized additive models

.

History

Though computational statistics is widely used today, it actually has a relatively short history of acceptance in the statistics community. For the most part, the founders of the field of statistics relied on mathematics and asymptotic approximations in the development of computational statistical methodology.^[4]

In statistical field, the first use of the term “computer” comes in an article in the Journal of the American Statistical Association archives by

electromechanical machine designed to assist in summarizing information stored on punched cards. It was invented by Herman Hollerith (February 29, 1860 – November 17, 1929), an American businessman, inventor, and statistician. His invention of the punched card tabulating machine was patented in 1884, and later was used in the 1890 Census of the United States. The advantages of the technology were immediately apparent. the 1880 Census, with about 50 million people, and it took over 7 years to tabulate. While in the 1890 Census, with over 62 million people, it took less than a year. This marks the beginning of the era of mechanized computational statistics and semiautomatic data processing

systems.

In 1908, William Sealy Gosset performed his now well-known Monte Carlo method simulation which led to the discovery of the Student’s t-distribution.^[5] With the help of computational methods, he also has plots of the empirical distributions overlaid on the corresponding theoretical distributions. The computer has revolutionized simulation and has made the replication of Gosset’s experiment little more than an exercise.^[6]^[7]

Later on, the scientists put forward computational ways of generating pseudo-random deviates, performed methods to convert uniform deviates into other distributional forms using inverse cumulative distribution function or acceptance-rejection methods, and developed state-space methodology for Markov chain Monte Carlo.^[8] One of the first efforts to generate random digits in a fully automated way, was undertaken by the RAND Corporation in 1947. The tables produced were published as a book in 1955, and also as a series of punch cards.

By the mid-1950s, several articles and patents for devices had been proposed for random number generators.^[9] The development of these devices were motivated from the need to use random digits to perform simulations and other fundamental components in statistical analysis. One of the most well known of such devices is ERNIE, which produces random numbers that determine the winners of the Premium Bond, a lottery bond issued in the United Kingdom. In 1958, John Tukey’s jackknife was developed. It is as a method to reduce the bias of parameter estimates in samples under nonstandard conditions.^[10] This requires computers for practical implementations. To this point, computers have made many tedious statistical studies feasible.^[11]

Methods

Maximum likelihood estimation

Maximum likelihood estimation is used to estimate the parameters of an assumed probability distribution, given some observed data. It is achieved by maximizing a likelihood function so that the observed data is most probable under the assumed statistical model.

Monte Carlo method

optimization, numerical integration, and generating draws from a probability distribution

.

Markov chain Monte Carlo

The

probability density proportional to a known function. These samples can be used to evaluate an integral over that variable, as its expected value or variance

. The more steps are included, the more closely the distribution of the sample matches the actual desired distribution.

Bootstrapping

The bootstrap is a resampling technique used to generate samples from an empirical probability distribution defined by an original sample of the population. It can be used to find a bootstrapped estimator of a population parameter. It can also be used to estimate the standard error of an estimator as well as to generate bootstrapped confidence intervals. The jackknife is a related technique^[12].

Applications

Computational statistics journals

Communications in Statistics - Simulation and Computation
Computational Statistics
Computational Statistics & Data Analysis
Journal of Computational and Graphical Statistics
Journal of Statistical Computation and Simulation
Journal of Statistical Software
The R Journal
The Stata Journal
Statistics and Computing
Wiley Interdisciplinary Reviews: Computational Statistics

Associations

International Association for Statistical Computing

References

^ Nolan, D. & Temple Lang, D. (2010). "Computing in the Statistics Curricula", The American Statistician 64 (2), pp.97-107.
^ ^a ^b Wegman, Edward J. “Computational Statistics: A New Agenda for Statistical Theory and Practice.” Journal of the Washington Academy of Sciences, vol. 78, no. 4, 1988, pp. 310–322. JSTOR
doi:10.1016/0167-9473(96)88920-1

S2CID 120111510
.

JSTOR 2331554.{{cite journal}}: CS1 maint: numeric names: authors list (link
)

OSTI 1569710. {{cite journal}}: Cite journal requires |journal= (help
)

PMID 18139350
.

S2CID 2806098
.

S2CID 4567651
.

ISSN 0006-3444
.

ISSN 0162-1459
.

ISBN 9781420010718
.

Further reading

Articles

Albert, J.H.; Gentle, J.E. (2004), Albert, James H; Gentle, James E (eds.), "Special Section: Teaching Computational Statistics", The American Statistician, 58: 1,
S2CID 219596225

Wilkinson, Leland (2008), "The Future of Statistical Computing (with discussion)", Technometrics, 50 (4): 418–435,
S2CID 3521989

Books

Drew, John H.;
ISBN 978-0-387-74675-3

Gentle, James E. (2002), Elements of Computational Statistics, Springer,
ISBN 0-387-95489-9

Gentle, James E.; Härdle, Wolfgang; Mori, Yuichi, eds. (2004), Handbook of Computational Statistics: Concepts and Methods, Springer,
ISBN 3-540-40464-3

Givens, Geof H.;
ISBN 978-0-471-46124-1

Klemens, Ben (2008), Modeling with Data: Tools and Techniques for Statistical Computing, Princeton University Press,
ISBN 978-0-691-13314-0

Monahan, John (2001), Numerical Methods of Statistics, Cambridge University Press,
ISBN 978-0-521-79168-7

Rose, Colin; Smith, Murray D. (2002), Mathematical Statistics with Mathematica, Springer Texts in Statistics, Springer,
ISBN 0-387-95234-9

Thisted, Ronald Aaron (1988), Elements of Statistical Computing: Numerical Computation, CRC Press,
ISBN 0-412-01371-1

Gharieb, Reda. R. (2017), Data Science: Scientific and Statistical Computing, Noor Publishing,
ISBN 978-3-330-97256-8

External links

Associations

International Association for Statistical Computing

Statistical Computing section of the American Statistical Association

Journals

Computational Statistics & Data Analysis

Journal of Computational & Graphical Statistics

Statistics and Computing

Authority control databases: National

Czech Republic

Retrieved from "https://en.wikipedia.org/w/index.php?title=Computational_statistics&oldid=1220008050"

[1] Nolan, D. & Temple Lang, D. (2010). "Computing in the Statistics Curricula", The American Statistician 64 (2), pp.97-107.

[:0-2] Wegman, Edward J. “Computational Statistics: A New Agenda for Statistical Theory and Practice.” Journal of the Washington Academy of Sciences, vol. 78, no. 4, 1988, pp. 310–322. JSTOR

[3] doi:10.1016/0167-9473(96)88920-1

[4] S2CID 120111510
.

[5] JSTOR 2331554.{{cite journal}}: CS1 maint: numeric names: authors list (link
)

[6] OSTI 1569710. {{cite journal}}: Cite journal requires |journal= (help
)

[7] PMID 18139350
.

[8] S2CID 2806098
.

[9] S2CID 4567651
.

[10] ISSN 0006-3444
.

[11] ISSN 0162-1459
.

[12] ISBN 9781420010718
.

[1]

[2]

[3]

[4]

[5]

[6]

[7]

[8]

[9]

[10]

[11]

[12]

History

Methods

Maximum likelihood estimation

Monte Carlo method

Markov chain Monte Carlo

Applications

Computational statistics journals

Associations

See also

References

Further reading

Articles

Books

External links

Associations

Journals