Ted Pedersen

odds.pm

# SYNOPSIS

Statistical library package to calculate the Odds Ratio. This package should be used with statistic.pl and rank.pl.

# DESCRIPTION

Assume that the frequency count data associated with a bigram <word1><word2> is stored in a 2x2 contingency table:

``````          word2   ~word2
word1    n11      n12 | n1p
~word1    n21      n22 | n2p
--------------
np1      np2   npp``````

where n11 is the number of times <word1><word2> occur together, and n12 is the number of times <word1> occurs with some word other than word2, and n1p is the number of times in total that word1 occurs as the first word in a bigram.

The odds ratio computes the ratio of the number of times that the words in a bigram occur together (or not at all) to the number of times the words occur individually. It is the cross product of the diagonal and the off-diagonal.

Thus, ODDS RATIO = n11*n22/n21*n12

if n21 and/or n12 is 0, then each zero value is "smoothed" to one to avoid a zero in the denominator.

# AUTHORS

Ted Pedersen <tpederse@d.umn.edu>

Bridget Thomson McInnes <bthomson@d.umn.edu>

# BUGS

This measure currently only defined for bigram data stored in 2x2 contingency table.