A comparison of two smoothing methods for word bigram models. Peto, L. B. Master's thesis, Department of Computer Science, University of Toronto, April, 1994. Published as technical report CSRI-304abstract bibtex Word bigram models estimated from text corpora require smoothing methods to estimate the probabilities of unseen bigrams. The deleted estimation method uses the formula:
Pr\(i|j\) = lambda fi + (1 - lambda)fi|j,
where fi and fi\|j are the relative frequency of i and the conditional relative frequency of i given j, respectively, and lambda is an optimized parameter. MacKay (1994) proposes a Bayesian approach using Dirichlet priors, which yields a different formula:
Pr\(i\|j\) = \(alpha/Fj \+ alpha\) mi \+ \(1 alpha/F_j \+ alpha\) fi\|j
where Fj is the count of j and alpha and mi are optimized parameters. This thesis describes an experiment in which the two methods were trained on a two-million-word corpus taken from the Canadian Hansard and compared on the basis of the experimental perplexity that they assigned to a shared test corpus. The methods proved to be about equally accurate, with MacKay's method using fewer resources.
@MastersThesis{ peto2,
author = {Linda Bauman Peto},
title = {A comparison of two smoothing methods for word bigram
models},
school = {Department of Computer Science, University of Toronto},
month = {April},
year = {1994},
note = {Published as technical report CSRI-304},
abstract = {<P>Word bigram models estimated from text corpora require
smoothing methods to estimate the probabilities of unseen
bigrams. The deleted estimation method uses the
formula:</p> <blockquote> Pr\(<I>i</I>|<I>j</I>\) =
<I>lambda</I> <I>f<SUB>i</SUB></I> + (1 -
<I>lambda</I>)<I>f<SUB>i|j</SUB></I>, </blockquote>
<P>where <I>f<SUB>i</SUB></I> and <I>f<SUB>i\|j</SUB></I>
are the relative frequency of <I>i</I> and the conditional
relative frequency of <I>i</I> given <I>j</I>,
respectively, and <I>lambda</I> is an optimized parameter.
MacKay (1994) proposes a Bayesian approach using Dirichlet
priors, which yields a different formula:</p> <BLOCKQUOTE>
Pr\(<I>i</I>\|<I>j</I>\) =
\(<I>alpha</I>/<I>F<SUB>j</SUB></I> \+ <I>alpha</I>\)
<I>m<SUB>i</SUB></I> \+ \(1 \- <I>alpha</I>/F_j \+
<I>alpha</I>\) <I>f<SUB>i\|j</SUB></I> </BLOCKQUOTE>
<P>where <I>F<SUB>j</SUB></I> is the count of <I>j</I> and
<I>alpha</I> and <I>m<SUB>i</SUB></I> are optimized
parameters. This thesis describes an experiment in which
the two methods were trained on a two-million-word corpus
taken from the Canadian <I>Hansard</I> and compared on the
basis of the experimental perplexity that they assigned to
a shared test corpus. The methods proved to be about
equally accurate, with MacKay's method using fewer
resources.</p>},
download = {http://ftp.cs.toronto.edu/pub/gh/Peto-MSc-1994.pdf}
}
Downloads: 0
{"_id":{"_str":"534282740e946d920a001ac8"},"__v":1,"authorIDs":["5538d3ec3affda41750016c9"],"author_short":["Peto, L. B."],"bibbaseid":"peto-acomparisonoftwosmoothingmethodsforwordbigrammodels-1994","bibdata":{"bibtype":"mastersthesis","type":"mastersthesis","author":[{"firstnames":["Linda","Bauman"],"propositions":[],"lastnames":["Peto"],"suffixes":[]}],"title":"A comparison of two smoothing methods for word bigram models","school":"Department of Computer Science, University of Toronto","month":"April","year":"1994","note":"Published as technical report CSRI-304","abstract":"<P>Word bigram models estimated from text corpora require smoothing methods to estimate the probabilities of unseen bigrams. The deleted estimation method uses the formula:</p> <blockquote> Pr\\(<I>i</I>|<I>j</I>\\) = <I>lambda</I> <I>f<SUB>i</SUB></I> + (1 - <I>lambda</I>)<I>f<SUB>i|j</SUB></I>, </blockquote> <P>where <I>f<SUB>i</SUB></I> and <I>f<SUB>i\\|j</SUB></I> are the relative frequency of <I>i</I> and the conditional relative frequency of <I>i</I> given <I>j</I>, respectively, and <I>lambda</I> is an optimized parameter. MacKay (1994) proposes a Bayesian approach using Dirichlet priors, which yields a different formula:</p> <BLOCKQUOTE> Pr\\(<I>i</I>\\|<I>j</I>\\) = \\(<I>alpha</I>/<I>F<SUB>j</SUB></I> \\+ <I>alpha</I>\\) <I>m<SUB>i</SUB></I> \\+ \\(1 <I>alpha</I>/F_j \\+ <I>alpha</I>\\) <I>f<SUB>i\\|j</SUB></I> </BLOCKQUOTE> <P>where <I>F<SUB>j</SUB></I> is the count of <I>j</I> and <I>alpha</I> and <I>m<SUB>i</SUB></I> are optimized parameters. This thesis describes an experiment in which the two methods were trained on a two-million-word corpus taken from the Canadian <I>Hansard</I> and compared on the basis of the experimental perplexity that they assigned to a shared test corpus. The methods proved to be about equally accurate, with MacKay's method using fewer resources.</p>","download":"http://ftp.cs.toronto.edu/pub/gh/Peto-MSc-1994.pdf","bibtex":"@MastersThesis{\t peto2,\n author\t= {Linda Bauman Peto},\n title\t\t= {A comparison of two smoothing methods for word bigram\n\t\t models},\n school\t= {Department of Computer Science, University of Toronto},\n month\t\t= {April},\n year\t\t= {1994},\n note\t\t= {Published as technical report CSRI-304},\n abstract\t= {<P>Word bigram models estimated from text corpora require\n\t\t smoothing methods to estimate the probabilities of unseen\n\t\t bigrams. The deleted estimation method uses the\n\t\t formula:</p> <blockquote> Pr\\(<I>i</I>|<I>j</I>\\) =\n\t\t <I>lambda</I> <I>f<SUB>i</SUB></I> + (1 -\n\t\t <I>lambda</I>)<I>f<SUB>i|j</SUB></I>, </blockquote>\n\t\t <P>where <I>f<SUB>i</SUB></I> and <I>f<SUB>i\\|j</SUB></I>\n\t\t are the relative frequency of <I>i</I> and the conditional\n\t\t relative frequency of <I>i</I> given <I>j</I>,\n\t\t respectively, and <I>lambda</I> is an optimized parameter.\n\t\t MacKay (1994) proposes a Bayesian approach using Dirichlet\n\t\t priors, which yields a different formula:</p> <BLOCKQUOTE>\n\t\t Pr\\(<I>i</I>\\|<I>j</I>\\) =\n\t\t \\(<I>alpha</I>/<I>F<SUB>j</SUB></I> \\+ <I>alpha</I>\\)\n\t\t <I>m<SUB>i</SUB></I> \\+ \\(1 \\- <I>alpha</I>/F_j \\+\n\t\t <I>alpha</I>\\) <I>f<SUB>i\\|j</SUB></I> </BLOCKQUOTE>\n\t\t <P>where <I>F<SUB>j</SUB></I> is the count of <I>j</I> and\n\t\t <I>alpha</I> and <I>m<SUB>i</SUB></I> are optimized\n\t\t parameters. This thesis describes an experiment in which\n\t\t the two methods were trained on a two-million-word corpus\n\t\t taken from the Canadian <I>Hansard</I> and compared on the\n\t\t basis of the experimental perplexity that they assigned to\n\t\t a shared test corpus. The methods proved to be about\n\t\t equally accurate, with MacKay's method using fewer\n\t\t resources.</p>},\n download\t= {http://ftp.cs.toronto.edu/pub/gh/Peto-MSc-1994.pdf}\n}\n\n","author_short":["Peto, L. B."],"key":"peto2","id":"peto2","bibbaseid":"peto-acomparisonoftwosmoothingmethodsforwordbigrammodels-1994","role":"author","urls":{},"metadata":{"authorlinks":{}}},"bibtype":"mastersthesis","biburl":"www.cs.toronto.edu/~fritz/tmp/compling.bib","downloads":0,"keywords":[],"search_terms":["comparison","two","smoothing","methods","word","bigram","models","peto"],"title":"A comparison of two smoothing methods for word bigram models","year":1994,"dataSources":["n8jB5BJxaeSmH6mtR","6b6A9kbkw4CsEGnRX"]}