Hidden Markov model

A hidden Markov model (HMM) is a

Markov process

(referred to as

X

). An HMM requires that there be an observable process

Y

whose outcomes depend on the outcomes of

X

in a known way. Since

X

cannot be observed directly, the goal is to learn about state of

X

by observing

Y

. By definition of being a Markov model, an HMM has an additional requirement that the outcome of

Y

at time

t=t_{0}

must be "influenced" exclusively by the outcome of

X

at

t=t_{0}

and that the outcomes of

X

and

Y

at

t<t_{0}

must be conditionally independent of

Y

at

t=t_{0}

given

X

at time

t=t_{0}

. Estimation of the parameters in an HMM can be performed using maximum likelihood estimation. For linear chain HMMs, the Baum–Welch algorithm can be used to estimate parameters.

Hidden Markov models are known for their applications to thermodynamics, statistical mechanics, physics, chemistry, economics, finance, signal processing, information theory, pattern recognition—such as speech,^[1] handwriting, gesture recognition,^[2] part-of-speech tagging, musical score following,^[3] partial discharges^[4] and bioinformatics.^[5]^[6]

Definition

Let $X_{n}$ and $Y_{n}$ be discrete-time stochastic processes and $n\geq 1$ . The pair $(X_{n},Y_{n})$ is a hidden Markov model if

$X_{n}$ is a
Markov process
whose behavior is not directly observable ("hidden");
$\operatorname {\mathbf {P} } {\bigl (}Y_{n}\in A\ {\bigl |}\ X_{1}=x_{1},\ldots ,X_{n}=x_{n}{\bigr )}=\operatorname {\mathbf {P} } {\bigl (}Y_{n}\in A\ {\bigl |}\ X_{n}=x_{n}{\bigr )}$ ,

for every

n\geq 1

,

x_{1},\ldots ,x_{n}

, and every Borel set

A

.

Let $X_{t}$ and $Y_{t}$ be continuous-time stochastic processes. The pair $(X_{t},Y_{t})$ is a hidden Markov model if

$X_{t}$ is a Markov process whose behavior is not directly observable ("hidden");
$\operatorname {\mathbf {P} } (Y_{t_{0}}\in A\mid \{X_{t}\in B_{t}\}_{t\leq t_{0}})=\operatorname {\mathbf {P} } (Y_{t_{0}}\in A\mid X_{t_{0}}\in B_{t_{0}})$ ,

for every

t_{0}

, every Borel set

A

, and every family of Borel sets

\{B_{t}\}_{t\leq t_{0}}

.

Terminology

The states of the process $X_{n}$ (resp. $X_{t})$ are called hidden states, and $\operatorname {\mathbf {P} } {\bigl (}Y_{n}\in A\mid X_{n}=x_{n}{\bigr )}$ (resp. $\operatorname {\mathbf {P} } {\bigl (}Y_{t}\in A\mid X_{t}\in B_{t}{\bigr )})$ is called emission probability or output probability.

Examples

Drawing balls from hidden urns

Figure 1. Probabilistic parameters of a hidden Markov model (example)
X — states
y — possible observations
a — state transition probabilities
b — output probabilities

In its discrete form, a hidden Markov process can be visualized as a generalization of the

Markov process

. It can be described by the upper part of Figure 1.

The Markov process cannot be observed, only the sequence of labeled balls, thus this arrangement is called a hidden Markov process. This is illustrated by the lower part of the diagram shown in Figure 1, where one can see that balls y1, y2, y3, y4 can be drawn at each state. Even if the observer knows the composition of the urns and has just observed a sequence of three balls, e.g. y1, y2 and y3 on the conveyor belt, the observer still cannot be sure which urn (i.e., at which state) the genie has drawn the third ball from. However, the observer can work out other information, such as the likelihood that the third ball came from each of the urns.

Weather guessing game

Consider two friends, Alice and Bob, who live far apart from each other and who talk together daily over the telephone about what they did that day. Bob is only interested in three activities: walking in the park, shopping, and cleaning his apartment. The choice of what to do is determined exclusively by the weather on a given day. Alice has no definite information about the weather, but she knows general trends. Based on what Bob tells her he did each day, Alice tries to guess what the weather must have been like.

Alice believes that the weather operates as a discrete Markov chain. There are two states, "Rainy" and "Sunny", but she cannot observe them directly, that is, they are hidden from her. On each day, there is a certain chance that Bob will perform one of the following activities, depending on the weather: "walk", "shop", or "clean". Since Bob tells Alice about his activities, those are the observations. The entire system is that of a hidden Markov model (HMM).

Alice knows the general weather trends in the area, and what Bob likes to do on average. In other words, the parameters of the HMM are known. They can be represented as follows in

Python

:

states = ("Rainy", "Sunny")

observations = ("walk", "shop", "clean")

start_probability = {"Rainy": 0.6, "Sunny": 0.4}

transition_probability = {
    "Rainy": {"Rainy": 0.7, "Sunny": 0.3},
    "Sunny": {"Rainy": 0.4, "Sunny": 0.6},
}

emission_probability = {
    "Rainy": {"walk": 0.1, "shop": 0.4, "clean": 0.5},
    "Sunny": {"walk": 0.6, "shop": 0.3, "clean": 0.1},
}

In this piece of code, start_probability represents Alice's belief about which state the HMM is in when Bob first calls her (all she knows is that it tends to be rainy on average). The particular probability distribution used here is not the equilibrium one, which is (given the transition probabilities) approximately {'Rainy': 0.57, 'Sunny': 0.43}. The transition_probability represents the change of the weather in the underlying Markov chain. In this example, there is only a 30% chance that tomorrow will be sunny if today is rainy. The emission_probability represents how likely Bob is to perform a certain activity on each day. If it is rainy, there is a 50% chance that he is cleaning his apartment; if it is sunny, there is a 60% chance that he is outside for a walk.

Graphical representation of the given HMM

A similar example is further elaborated in the Viterbi algorithm page.

Structural architecture

The diagram below shows the general architecture of an instantiated HMM. Each oval shape represents a random variable that can adopt any of a number of values. The random variable x(t) is the hidden state at time $t$ (with the model from the above diagram, x(t) ∈ { x₁, x₂, x₃ }). The random variable y(t) is the observation at time $t$ (with y(t) ∈ { y₁, y₂, y₃, y₄ }). The arrows in the diagram (often called a trellis diagram) denote conditional dependencies.

From the diagram, it is clear that the conditional probability distribution of the hidden variable x(t) at time $t$ , given the values of the hidden variable $x$ at all times, depends only on the value of the hidden variable x(t − 1); the values at time t − 2 and before have no influence. This is called the Markov property. Similarly, the value of the observed variable y(t) depends on only the value of the hidden variable x(t) (both at time $t$ ).

In the standard type of hidden Markov model considered here, the state space of the hidden variables is discrete, while the observations themselves can either be discrete (typically generated from a

Gaussian distribution

). The parameters of a hidden Markov model are of two types, transition probabilities and emission probabilities (also known as output probabilities). The transition probabilities control the way the hidden state at time

t

is chosen given the hidden state at time

t-1

.

The hidden state space is assumed to consist of one of $N$ possible values, modelled as a categorical distribution. (See the section below on extensions for other possibilities.) This means that for each of the $N$ possible states that a hidden variable at time $t$ can be in, there is a transition probability from this state to each of the $N$ possible states of the hidden variable at time $t+1$ , for a total of $N^{2}$ transition probabilities. The set of transition probabilities for transitions from any given state must sum to 1. Thus, the $N\times N$ matrix of transition probabilities is a Markov matrix. Because any transition probability can be determined once the others are known, there are a total of $N(N-1)$ transition parameters.

In addition, for each of the $N$ possible states, there is a set of emission probabilities governing the distribution of the observed variable at a particular time given the state of the hidden variable at that time. The size of this set depends on the nature of the observed variable. For example, if the observed variable is discrete with $M$ possible values, governed by a categorical distribution, there will be $M-1$ separate parameters, for a total of $N(M-1)$ emission parameters over all hidden states. On the other hand, if the observed variable is an $M$ -dimensional vector distributed according to an arbitrary

multivariate Gaussian distribution, there will be

M

parameters controlling the means

and

{\frac {M(M+1)}{2}}

parameters controlling the covariance matrix, for a total of

N\left(M+{\frac {M(M+1)}{2}}\right)={\frac {NM(M+3)}{2}}=O(NM^{2})

emission parameters. (In such a case, unless the value of

M

is small, it may be more practical to restrict the nature of the covariances between individual elements of the observation vector, e.g. by assuming that the elements are independent of each other, or less restrictively, are independent of all but a fixed number of adjacent elements.)

Inference

The state transition and output probabilities of an HMM are indicated by the line opacity in the upper part of the diagram. Given that the output sequence is observed in the lower part of the diagram, interest occurs in the most likely sequence of states that could have produced it. Based on the arrows that are present in the diagram, the following state sequences are candidates:
5 3 2 5 3 2
4 3 2 5 3 2
3 1 2 5 3 2
The most likely sequence can be found by evaluating the joint probability of both the state sequence and the observations for each case (simply by multiplying the probability values, which here correspond to the opacities of the arrows involved). In general, this type of problem (i.e., finding the most likely explanation for an observation sequence) can be solved efficiently using the Viterbi algorithm.

Several inference problems are associated with hidden Markov models, as outlined below.

Probability of an observed sequence

The task is to compute in a best way, given the parameters of the model, the probability of a particular output sequence. This requires summation over all possible state sequences:

The probability of observing a sequence

Y=y(0),y(1),\dots ,y(L-1),

of length L is given by

P(Y)=\sum _{X}P(Y\mid X)P(X),

where the sum runs over all possible hidden-node sequences

X=x(0),x(1),\dots ,x(L-1).

Applying the principle of dynamic programming, this problem, too, can be handled efficiently using the forward algorithm.

Probability of the latent variables

A number of related tasks ask about the probability of one or more of the latent variables, given the model's parameters and a sequence of observations $y(1),\dots ,y(t)$ .

Filtering

The task is to compute, given the model's parameters and a sequence of observations, the distribution over hidden states of the last latent variable at the end of the sequence, i.e. to compute $P(x(t)\mid y(1),\dots ,y(t))$ . This task is used when the sequence of latent variables is thought of as the underlying states that a process moves through at a sequence of points in time, with corresponding observations at each point. Then, it is natural to ask about the state of the process at the end.

This problem can be handled efficiently using the forward algorithm. An example is when the algorithm is applied to a Hidden Markov Network to determine $\mathrm {P} {\big (}h_{t}\mid v_{1:t}{\big )}$ .

Smoothing

This is similar to filtering but asks about the distribution of a latent variable somewhere in the middle of a sequence, i.e. to compute $P(x(k)\mid y(1),\dots ,y(t))$ for some $k<t$ . From the perspective described above, this can be thought of as the probability distribution over hidden states for a point in time k in the past, relative to time t.

The

forward-backward algorithm

is a good method for computing the smoothed values for all hidden state variables.

Most likely explanation

The task, unlike the previous two, asks about the

joint probability of the entire sequence of hidden states that generated a particular sequence of observations (see illustration on the right). This task is generally applicable when HMM's are applied to different sorts of problems from those for which the tasks of filtering and smoothing are applicable. An example is part-of-speech tagging, where the hidden states represent the underlying parts of speech

corresponding to an observed sequence of words. In this case, what is of interest is the entire sequence of parts of speech, rather than simply the part of speech for a single word, as filtering or smoothing would compute.

This task requires finding a maximum over all possible state sequences, and can be solved efficiently by the Viterbi algorithm.

Statistical significance

For some of the above problems, it may also be interesting to ask about statistical significance. What is the probability that a sequence drawn from some null distribution will have an HMM probability (in the case of the forward algorithm) or a maximum state sequence probability (in the case of the Viterbi algorithm) at least as large as that of a particular output sequence?^[8] When an HMM is used to evaluate the relevance of a hypothesis for a particular output sequence, the statistical significance indicates the false positive rate associated with failing to reject the hypothesis for the output sequence.

Learning

The parameter learning task in HMMs is to find, given an output sequence or a set of such sequences, the best set of state transition and emission probabilities. The task is usually to derive the

expectation-maximization algorithm

.

If the HMMs are used for time series prediction, more sophisticated Bayesian inference methods, like Markov chain Monte Carlo (MCMC) sampling are proven to be favorable over finding a single maximum likelihood model both in terms of accuracy and stability.^[9] Since MCMC imposes significant computational burden, in cases where computational scalability is also of interest, one may alternatively resort to variational approximations to Bayesian inference, e.g.^[10] Indeed, approximate variational inference offers computational efficiency comparable to expectation-maximization, while yielding an accuracy profile only slightly inferior to exact MCMC-type Bayesian inference.

Applications

HMMs can be applied in many fields where the goal is to recover a data sequence that is not immediately observable (but other data that depend on the sequence are). Applications include:

Computational finance^[11]^[12]
Single-molecule kinetic analysis^[13]
Neuroscience^[14]^[15]
Cryptanalysis
Speech recognition, including Siri^[16]
Speech synthesis
Part-of-speech tagging
Document separation in scanning solutions
Machine translation
Partial discharge
Gene prediction
Handwriting recognition^[17]
Alignment of bio-sequences
Time series analysis
Activity recognition
Protein folding^[18]
Sequence classification^[19]
Metamorphic virus detection^[20]
Sequence motif discovery (DNA and proteins)^[21]
DNA hybridization kinetics^[22]^[23]
Chromatin state discovery^[24]
Transportation forecasting^[25]
Solar irradiance variability^[26]^[27]^[28]

History

Hidden Markov models were described in a series of statistical papers by Leonard E. Baum and other authors in the second half of the 1960s.^[29]^[30]^[31]^[32]^[33] One of the first applications of HMMs was speech recognition, starting in the mid-1970s.^[34]^[35]^[36]^[37] From the linguistics point of view, hidden Markov models are equivalent to stochastic regular grammar.^[38]

In the second half of the 1980s, HMMs began to be applied to the analysis of biological sequences,^[39] in particular DNA. Since then, they have become ubiquitous in the field of bioinformatics.^[40]

Extensions

General state spaces

In the hidden Markov models considered above, the state space of the hidden variables is discrete, while the observations themselves can either be discrete (typically generated from a

Gaussian distribution. In simple cases, such as the linear dynamical system just mentioned, exact inference is tractable (in this case, using the Kalman filter); however, in general, exact inference in HMMs with continuous latent variables is infeasible, and approximate methods must be used, such as the extended Kalman filter or the particle filter

.

Nowadays, inference in hidden Markov models is performed in nonparametric settings, where the dependency structure enables identifiability of the model^[41] and the learnability limits are still under exploration.^[42]

Bayesian modeling of the transitions probabilities

Hidden Markov models are

expectation-maximization algorithm

.

An extension of the previously described hidden Markov models with Dirichlet priors uses a Dirichlet process in place of a Dirichlet distribution. This type of model allows for an unknown and potentially infinite number of states. It is common to use a two-level Dirichlet process, similar to the previously described model with two levels of Dirichlet distributions. Such a model is called a hierarchical Dirichlet process hidden Markov model, or HDP-HMM for short. It was originally described under the name "Infinite Hidden Markov Model"^[43] and was further formalized in "Hierarchical Dirichlet Processes".^[44]

Discriminative approach

A different type of extension uses a

statistically independent

of each other, as would be the case if such features were used in a generative model. Finally, arbitrary features over pairs of adjacent hidden states can be used rather than simple transition probabilities. The disadvantages of such models are: (1) The types of prior distributions that can be placed on hidden states are severely limited; (2) It is not possible to predict the probability of seeing an arbitrary observation. This second limitation is often not an issue in practice, since many common usages of HMM's do not require such predictive probabilities.

A variant of the previously described discriminative model is the linear-chain conditional random field. This uses an undirected graphical model (aka Markov random field) rather than the directed graphical models of MEMM's and similar models. The advantage of this type of model is that it does not suffer from the so-called label bias problem of MEMM's, and thus may make more accurate predictions. The disadvantage is that training can be slower than for MEMM's.

Other extensions

Yet another variant is the factorial hidden Markov model, which allows for a single observation to be conditioned on the corresponding hidden variables of a set of $K$ independent Markov chains, rather than a single Markov chain. It is equivalent to a single HMM, with $N^{K}$ states (assuming there are $N$ states for each chain), and therefore, learning in such a model is difficult: for a sequence of length $T$ , a straightforward Viterbi algorithm has complexity $O(N^{2K}\,T)$ . To find an exact solution, a junction tree algorithm could be used, but it results in an $O(N^{K+1}\,K\,T)$ complexity. In practice, approximate techniques, such as variational approaches, could be used.^[45]

All of the above models can be extended to allow for more distant dependencies among hidden states, e.g. allowing for a given state to be dependent on the previous two or three states rather than a single previous state; i.e. the transition probabilities are extended to encompass sets of three or four adjacent states (or in general $K$ adjacent states). The disadvantage of such models is that dynamic-programming algorithms for training them have an $O(N^{K}\,T)$ running time, for $K$ adjacent states and $T$ total observations (i.e. a length- $T$ Markov chain). This extension has been widely used in bioinformatics, in the modeling of DNA sequences.

Another recent extension is the triplet Markov model,^[46] in which an auxiliary underlying process is added to model some data specificities. Many variants of this model have been proposed. One should also mention the interesting link that has been established between the theory of evidence and the triplet Markov models^[47] and which allows to fuse data in Markovian context^[48] and to model nonstationary data.^[49]^[50] Alternative multi-stream data fusion strategies have also been proposed in recent literature, e.g.,^[51]

Finally, a different rationale towards addressing the problem of modeling nonstationary data by means of hidden Markov models was suggested in 2012.^[52] It consists in employing a small recurrent neural network (RNN), specifically a reservoir network,^[53] to capture the evolution of the temporal dynamics in the observed data. This information, encoded in the form of a high-dimensional vector, is used as a conditioning variable of the HMM state transition probabilities. Under such a setup, eventually is obtained a nonstationary HMM, the transition probabilities of which evolve over time in a manner that is inferred from the data, in contrast to some unrealistic ad-hoc model of temporal evolution.

In 2023, two innovative algorithms were introduced for the Hidden Markov Model. These algorithms enable the computation of the posterior distribution of the HMM without the necessity of explicitly modeling the joint distribution, utilizing only the conditional distributions.^[54]^[55] Unlike traditional methods such as the Forward-Backward and Viterbi algorithms, which require knowledge of the joint law of the HMM and can be computationally intensive to learn, the Discriminative Forward-Backward and Discriminative Viterbi algorithms circumvent the need for the observation's law.^[56]^[57] This breakthrough allows the HMM to be applied as a discriminative model, offering a more efficient and versatile approach to leveraging Hidden Markov Models in various applications.

The model suitable in the context of longitudinal data is named latent Markov model.^[58] The basic version of this model has been extended to include individual covariates, random effects and to model more complex data structures such as multilevel data. A complete overview of the latent Markov models, with special attention to the model assumptions and to their practical use is provided in^[59]

Measure theory

Given a Markov transition matrix and an invariant distribution on the states, a probability measure can be imposed on the set of subshifts. For example, consider the Markov chain given on the left on the states $A,B_{1},B_{2}$ , with invariant distribution $\pi =(2/7,4/7,1/7)$ . By ignoring the distinction between $B_{1},B_{2}$ , this space of subshifts is projected on $A,B_{1},B_{2}$ into another space of subshifts on $A,B$ , and this projection also projects the probability measure down to a probability measure on the subshifts on $A,B$ .

The curious thing is that the probability measure on the subshifts on $A,B$ is not created by a Markov chain on $A,B$ , not even multiple orders. Intuitively, this is because if one observes a long sequence of $B^{n}$ , then one would become increasingly sure that the $\Pr(A\mid B^{n})\to {\frac {2}{3}}$ , meaning that the observable part of the system can be affected by something infinitely in the past.^[60]^[61]

Conversely, there exists a space of subshifts on 6 symbols, projected to subshifts on 2 symbols, such that any Markov measure on the smaller subshift has a preimage measure that is not Markov of any order (example 2.6^[61]).

References

^ "Google Scholar".
^ Thad Starner, Alex Pentland. Real-Time American Sign Language Visual Recognition From Video Using Hidden Markov Models. Master's Thesis, MIT, Feb 1995, Program in Media Arts
^ B. Pardo and W. Birmingham. Modeling Form for On-line Following of Musical Performances Archived 2012-02-06 at the Wayback Machine. AAAI-05 Proc., July 2005.
^ Satish L, Gururaj BI (April 2003). "Use of hidden Markov models for partial discharge pattern classification". IEEE Transactions on Dielectrics and Electrical Insulation.
PMID 14704198
.

PMID 22373907
.

S2CID 13618539. [1]

PMID 19589158.

^ Sipos, I. Róbert. Parallel stratified MCMC sampling of AR-HMMs for stochastic time series prediction. In: Proceedings, 4th Stochastic Modeling Techniques and Data Analysis International Conference with Demographics Workshop (SMTDA2016), pp. 295-306. Valletta, 2016. PDF

doi:10.1016/j.patcog.2010.09.001. Archived from the original
(PDF) on 2011-04-01. Retrieved 2018-03-11.

S2CID 61882456
.

doi:10.1016/j.eswa.2016.01.015
.

doi:10.1142/S1793048013300053
.

PMID 35302683
.

S2CID 235703641
.

ISBN 9780465061921
.

^ Kundu, Amlan, Yang He, and Paramvir Bahl. "Recognition of handwritten word: first and second order hidden Markov model based approach^{[dead link]}." Pattern recognition 22.3 (1989): 283-297.

S2CID 5502662
.

^ Blasiak, S.; Rangwala, H. (2011). "A Hidden Markov Model Variant for Sequence Classification". IJCAI Proceedings-International Joint Conference on Artificial Intelligence. 22: 1192.

S2CID 8116065
.

PMID 23814189
.

S2CID 96448257
.

S2CID 84841635
.

^ "ChromHMM: Chromatin state discovery and characterization". compbio.mit.edu. Retrieved 2018-08-01.

arXiv:1707.09133 [stat.AP
].

doi:10.1016/S0038-092X(98)00004-8
.

S2CID 125867684
.

S2CID 125538244
.

doi:10.1214/aoms/1177699147
.

Zbl 0157.11101
.

doi:10.2140/pjm.1968.27.211
.

Zbl 0188.49603
.

^ Baum, L.E. (1972). "An Inequality and Associated Maximization Technique in Statistical Estimation of Probabilistic Functions of a Markov Process". Inequalities. 3: 1–8.

doi:10.1109/TASSP.1975.1162650
.

doi:10.1109/TIT.1975.1055384
.

ISBN 978-0-7486-0162-2
.

ISBN 978-0-13-022616-7
.

ISBN 978-3-540-48985-6
.

PMID 3641921. (subscription required)

OCLC 593254083

ISSN 1573-1375
.

ISSN 0018-9448
.

^ Beal, Matthew J., Zoubin Ghahramani, and Carl Edward Rasmussen. "The infinite hidden Markov model." Advances in neural information processing systems 14 (2002): 577-584.

^ Teh, Yee Whye, et al. "Hierarchical dirichlet processes." Journal of the American Statistical Association 101.476 (2006).

doi:10.1023/A:1007425814087
.

doi:10.1016/S1631-073X(02)02462-7
.

doi:10.1016/j.ijar.2006.05.001
.

^ Boudaren et al. Archived 2014-03-11 at the Wayback Machine, M. Y. Boudaren, E. Monfrini, W. Pieczynski, and A. Aissani, Dempster-Shafer fusion of multisensor signals in nonstationary Markovian context, EURASIP Journal on Advances in Signal Processing, No. 134, 2012.

^ Lanchantin et al., P. Lanchantin and W. Pieczynski, Unsupervised restoration of hidden non stationary Markov chain using evidential priors, IEEE Transactions on Signal Processing, Vol. 53, No. 8, pp. 3091-3098, 2005.

^ Boudaren et al., M. Y. Boudaren, E. Monfrini, and W. Pieczynski, Unsupervised segmentation of random discrete data hidden with switching noise distributions, IEEE Signal Processing Letters, Vol. 19, No. 10, pp. 619-622, October 2012.

^ Sotirios P. Chatzis, Dimitrios Kosmopoulos, "Visual Workflow Recognition Using a Variational Bayesian Treatment of Multistream Fused Hidden Markov Models," IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, no. 7, pp. 1076-1086, July 2012.

hdl:10044/1/12611
.

^ M. Lukosevicius, H. Jaeger (2009) Reservoir computing approaches to recurrent neural network training, Computer Science Review 3: 127–149.

^ Azeraf, E., Monfrini, E., & Pieczynski, W. (2023). Equivalence between LC-CRF and HMM, and Discriminative Computing of HMM-Based MPM and MAP. Algorithms, 16(3), 173.

^ Azeraf, E., Monfrini, E., Vignon, E., & Pieczynski, W. (2020). Hidden markov chains, entropic forward-backward, and part-of-speech tagging. arXiv preprint arXiv:2005.10629.

^ Azeraf, E., Monfrini, E., & Pieczynski, W. (2022). Deriving discriminative classifiers from generative models. arXiv preprint arXiv:2201.00844.

^ Ng, A., & Jordan, M. (2001). On discriminative vs. generative classifiers: A comparison of logistic regression and naive bayes. Advances in neural information processing systems, 14.

^ Wiggins, L. M. (1973). Panel Analysis: Latent Probability Models for Attitude and Behaviour Processes. Amsterdam: Elsevier.

ISBN 978-14-3981-708-7
.

^ Sofic Measures: Characterizations of Hidden Markov Chains by Linear Algebra, Formal Languages, and Symbolic Dynamics - Karl Petersen, Mathematics 210, Spring 2006, University of North Carolina at Chapel Hill

^
arXiv:0907.1858

External links

Wikimedia Commons has media related to Hidden Markov Model.

Concepts

Teif, V. B.; Rippe, K. (2010). "Statistical–mechanical lattice models for protein–DNA binding in chromatin". J. Phys.: Condens. Matter. 22 (41): 414105.
S2CID 103345
.

A Revealing Introduction to Hidden Markov Models by Mark Stamp, San Jose State University.

Fitting HMM's with expectation-maximization – complete derivation

A step-by-step tutorial on HMMs Archived 2017-08-13 at the Wayback Machine (University of Leeds)

Hidden Markov Models (an exposition using basic mathematics)

Hidden Markov Models (by Narada Warakagoda)

Hidden Markov Models: Fundamentals and Applications Part 1, Part 2 (by V. Petrushin)

Lecture on a Spreadsheet by Jason Eisner, Video and interactive spreadsheet

Discrete time

Bernoulli process

Branching process

Chinese restaurant process

Galton–Watson process

Independent and identically distributed random variables

Markov chain

Moran process

Random walk
Loop-erased

Self-avoiding

Biased

Maximal entropy

Continuous time

Additive process

Bessel process

Birth–death process
pure birth

Brownian motion
Bridge

Excursion

Fractional

Geometric

Meander

Cauchy process

Contact process

Continuous-time random walk

Cox process

Diffusion process

Dyson Brownian motion

Empirical process

Feller process

Fleming–Viot process

Gamma process

Geometric process

Hawkes process

Hunt process

Interacting particle systems

Itô diffusion

Itô process

Jump diffusion

Jump process

Lévy process

Local time

Markov additive process

McKean–Vlasov process

Ornstein–Uhlenbeck process

Poisson process
Compound

Non-homogeneous

Quasimartingale

Schramm–Loewner evolution

Semimartingale

Sigma-martingale

Stable process

Superprocess

Telegraph process

Variance gamma process

Wiener process

Wiener sausage

Both

Branching process

Gaussian process

Hidden Markov model (HMM)

Markov process

Martingale
Differences

Local

Sub-

Super-

Random dynamical system

Regenerative process

Renewal process

Stochastic chains with memory of variable length

White noise

Fields and other

Dirichlet process

Gaussian random field

Gibbs measure

Hopfield model

Ising model
Potts model

Boolean network

Markov random field

Percolation

Pitman–Yor process

Point process
Cox

Poisson

Random field

Random graph

Time series models

Autoregressive conditional heteroskedasticity (ARCH) model

Autoregressive integrated moving average (ARIMA) model

Autoregressive (AR) model

Autoregressive–moving-average (ARMA) model

Generalized autoregressive conditional heteroskedasticity (GARCH) model

Moving-average (MA) model

Financial models

Binomial options pricing model

Black–Derman–Toy

Black–Karasinski

Black–Scholes

Chan–Karolyi–Longstaff–Sanders (CKLS)

Chen

Constant elasticity of variance (CEV)

Cox–Ingersoll–Ross (CIR)

Garman–Kohlhagen

Heath–Jarrow–Morton (HJM)

Heston

Ho–Lee

Hull–White

Korn-Kreer-Lenssen

LIBOR market

Rendleman–Bartter

SABR volatility

Vašíček

Wilkie

Actuarial models

Bühlmann

Cramér–Lundberg

Risk process

Sparre–Anderson

Queueing models

Bulk

Fluid

Generalized queueing network

M/G/1

M/M/1

M/M/c

Properties

Càdlàg paths

Continuous

Continuous paths

Ergodic

Exchangeable

Feller-continuous

Gauss–Markov

Markov

Mixing

Piecewise-deterministic

Predictable

Progressively measurable

Self-similar

Stationary

Time-reversible

Limit theorems

Central limit theorem

Donsker's theorem

Doob's martingale convergence theorems

Ergodic theorem

Fisher–Tippett–Gnedenko theorem

Large deviation principle

Law of large numbers (weak/strong)

Law of the iterated logarithm

Maximal ergodic theorem

Sanov's theorem

Lévy
)

Inequalities

Burkholder–Davis–Gundy

Doob's martingale

Doob's upcrossing

Kunita–Watanabe

Marcinkiewicz–Zygmund

Tools

Cameron–Martin formula

Convergence of random variables

Doléans-Dade exponential

Doob decomposition theorem

Doob–Meyer decomposition theorem

Doob's optional stopping theorem

Dynkin's formula

Feynman–Kac formula

Filtration

Girsanov theorem

Infinitesimal generator

Itô integral

Itô's lemma

Karhunen–Loève theorem

Kolmogorov continuity theorem

Kolmogorov extension theorem

Lévy–Prokhorov metric

Malliavin calculus

Martingale representation theorem

Optional stopping theorem

Prokhorov's theorem

Quadratic variation

Reflection principle

Skorokhod integral

Skorokhod's representation theorem

Skorokhod space

Snell envelope

Stochastic differential equation
Tanaka

Stopping time

Stratonovich integral

Uniform integrability

Usual hypotheses

Wiener space

Classical

Abstract

Disciplines

Actuarial mathematics

Control theory

Econometrics

Ergodic theory

Extreme value theory (EVT)

Large deviations theory

Mathematical finance

Mathematical statistics

Probability theory

Queueing theory

Renewal theory

Ruin theory

Signal processing

Statistics

Stochastic analysis

Time series analysis

Machine learning

List of topics

Category

Authority control databases: National
Germany
United States
Israel

Retrieved from "https://en.wikipedia.org/w/index.php?title=Hidden_Markov_model&oldid=1295082173"

[1] "Google Scholar".

[2] Thad Starner, Alex Pentland. Real-Time American Sign Language Visual Recognition From Video Using Hidden Markov Models. Master's Thesis, MIT, Feb 1995, Program in Media Arts

[3] B. Pardo and W. Birmingham. Modeling Form for On-line Following of Musical Performances Archived 2012-02-06 at the Wayback Machine. AAAI-05 Proc., July 2005.

[4] Satish L, Gururaj BI (April 2003). "Use of hidden Markov models for partial discharge pattern classification". IEEE Transactions on Dielectrics and Electrical Insulation.

[5] PMID 14704198
.

[6] PMID 22373907
.

[7] S2CID 13618539. [1]

[8] PMID 19589158.

[9] Sipos, I. Róbert. Parallel stratified MCMC sampling of AR-HMMs for stochastic time series prediction. In: Proceedings, 4th Stochastic Modeling Techniques and Data Analysis International Conference with Demographics Workshop (SMTDA2016), pp. 295-306. Valletta, 2016. PDF

[10] :10.1016/j.patcog.2010.09.001. Archived from the original
(PDF) on 2011-04-01. Retrieved 2018-03-11.

[11] S2CID 61882456
.

[12] :10.1016/j.eswa.2016.01.015
.

[13] :10.1142/S1793048013300053
.

[14] PMID 35302683
.

[15] S2CID 235703641
.

[16] ISBN 9780465061921
.

[17] Kundu, Amlan, Yang He, and Paramvir Bahl. "Recognition of handwritten word: first and second order hidden Markov model based approach^{[dead link]}." Pattern recognition 22.3 (1989): 283-297.

[18] S2CID 5502662
.

[19] Blasiak, S.; Rangwala, H. (2011). "A Hidden Markov Model Variant for Sequence Classification". IJCAI Proceedings-International Joint Conference on Artificial Intelligence. 22: 1192.

[20] S2CID 8116065
.

[21] PMID 23814189
.

[22] S2CID 96448257
.

[23] S2CID 84841635
.

[24] "ChromHMM: Chromatin state discovery and characterization". compbio.mit.edu. Retrieved 2018-08-01.

[25] rXiv:1707.09133 [stat.AP
].

[26] :10.1016/S0038-092X(98)00004-8
.

[27] S2CID 125867684
.

[28] S2CID 125538244
.

[29] :10.1214/aoms/1177699147
.

[30] Zbl 0157.11101
.

[31] :10.2140/pjm.1968.27.211
.

[32] Zbl 0188.49603
.

[33] Baum, L.E. (1972). "An Inequality and Associated Maximization Technique in Statistical Estimation of Probabilistic Functions of a Markov Process". Inequalities. 3: 1–8.

[34] :10.1109/TASSP.1975.1162650
.

[35] :10.1109/TIT.1975.1055384
.

[36] ISBN 978-0-7486-0162-2
.

[37] ISBN 978-0-13-022616-7
.

[38] ISBN 978-3-540-48985-6
.

[39] PMID 3641921. (subscription required)

[durbin-40] OCLC 593254083

[41] ISSN 1573-1375
.

[42] ISSN 0018-9448
.

[43] Beal, Matthew J., Zoubin Ghahramani, and Carl Edward Rasmussen. "The infinite hidden Markov model." Advances in neural information processing systems 14 (2002): 577-584.

[44] Teh, Yee Whye, et al. "Hierarchical dirichlet processes." Journal of the American Statistical Association 101.476 (2006).

[45] doi:10.1023/A:1007425814087
.

[TMM-46] :10.1016/S1631-073X(02)02462-7
.

[TMMEV-47] :10.1016/j.ijar.2006.05.001
.

[JASP-48] Boudaren et al. Archived 2014-03-11 at the Wayback Machine, M. Y. Boudaren, E. Monfrini, W. Pieczynski, and A. Aissani, Dempster-Shafer fusion of multisensor signals in nonstationary Markovian context, EURASIP Journal on Advances in Signal Processing, No. 134, 2012.

[TSP-49] Lanchantin et al., P. Lanchantin and W. Pieczynski, Unsupervised restoration of hidden non stationary Markov chain using evidential priors, IEEE Transactions on Signal Processing, Vol. 53, No. 8, pp. 3091-3098, 2005.

[SPL-50] Boudaren et al., M. Y. Boudaren, E. Monfrini, and W. Pieczynski, Unsupervised segmentation of random discrete data hidden with switching noise distributions, IEEE Signal Processing Letters, Vol. 19, No. 10, pp. 619-622, October 2012.

[51] Sotirios P. Chatzis, Dimitrios Kosmopoulos, "Visual Workflow Recognition Using a Variational Bayesian Treatment of Multistream Fused Hidden Markov Models," IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, no. 7, pp. 1076-1086, July 2012.

[Reservoir-HMM-52] hdl:10044/1/12611
.

[53] M. Lukosevicius, H. Jaeger (2009) Reservoir computing approaches to recurrent neural network training, Computer Science Review 3: 127–149.

[54] Azeraf, E., Monfrini, E., & Pieczynski, W. (2023). Equivalence between LC-CRF and HMM, and Discriminative Computing of HMM-Based MPM and MAP. Algorithms, 16(3), 173.

[55] Azeraf, E., Monfrini, E., Vignon, E., & Pieczynski, W. (2020). Hidden markov chains, entropic forward-backward, and part-of-speech tagging. arXiv preprint arXiv:2005.10629.

[56] Azeraf, E., Monfrini, E., & Pieczynski, W. (2022). Deriving discriminative classifiers from generative models. arXiv preprint arXiv:2201.00844.

[57] Ng, A., & Jordan, M. (2001). On discriminative vs. generative classifiers: A comparison of logistic regression and naive bayes. Advances in neural information processing systems, 14.

[58] Wiggins, L. M. (1973). Panel Analysis: Latent Probability Models for Attitude and Behaviour Processes. Amsterdam: Elsevier.

[59] ISBN 978-14-3981-708-7
.

[:0-60] Sofic Measures: Characterizations of Hidden Markov Chains by Linear Algebra, Formal Languages, and Symbolic Dynamics - Karl Petersen, Mathematics 210, Spring 2006, University of North Carolina at Chapel Hill

[:1-61] 
arXiv:0907.1858

[1]

[2]

[3]

[4]

[5]

[6]

[8]

[9]

[10]

[11]

[12]

[13]

[14]

[15]

[16]

[17]

[18]

[19]

[20]

[21]

[22]

[23]

[24]

[25]

[26]

[27]

[28]

[29]

[30]

[31]

[32]

[33]

[34]

[35]

[36]

[37]

[38]

[39]

[40]

[41]

[42]

[43]

[44]

[45]

[46]

[47]

[48]

[49]

[50]

[51]

[52]

[53]

[54]

[55]

[56]

[57]

[58]

[59]

[60]

[61]