The Wayback Machine - https://web.archive.org/web/20110426125750/http://citeseerx.ist.psu.edu:80/viewdoc/summary?doi=10.1.1.14.5962
MetaCart Sign in to MyCiteSeerX

Accurate Methods for the Statistics of Surprise and Coincidence (1993) [555 citations — 1 self]

Abstract:

Much work has been done on the statistical analysis of text. In some cases reported in the literature, inappropriate statistical methods have been used, and statistical significance of results have not been addressed. In particular, asymptotic normality assumptions have often been used unjustifiably, leading to flawed results.This assumption of normal distribution limits the ability to analyze rare events. Unfortunately rare events do make up a large fraction of real text.However, more applicable methods based on likelihood ratio tests are available that yield good results with relatively small samples. These tests can be implemented efficiently, and have been used for the detection of composite terms and for the determination of domain-specific terms. In some cases, these measures perform much better than the methods previously used. In cases where traditional contingency table methods work well, the likelihood ratio tests described here are nearly identical.This paper describes the basis of a measure based on likelihood ratios that can be applied to the analysis of text.

Citations

439 A statistical approach to machine translation – Brown, Cocke, et al. - 1990
134 Introduction to the Theory of Statistics – Mood, Graybill, et al. - 1974
131 Identifying word correspondence in parallel texts – Gale, Church - 1991
77 Using latent semantic analysis to improve access to textual information – Dumais, Furnas, et al. - 1988
27 Parsing, Word Associations and Typical Predicate-Argument Relations – Church - 1989
24 Distribution-free statistical tests – Bradley - 1968
1 kji Pji - ~j kji = ~kjlogpj J – Gale, Church - 1993