Tuesday, May 19, 2009

How to extract words for dictionary - 사전 만드는 방법

Corpus-Specific Stemming gathers unique words from the corpus, makes equivalence classes, and after some statistical calculations and reclassification makes a dictionary. [1]

References
[ 1 ] Croft, W. B. and Xu, J. “Corpus-Specific Stemming using Word Form Co-occurrences”. In Fourth Annual Symposium on Document Analysis and Information Retrieval, 1995.
[ 2 ] Krovetz, R. “View Morphology as an Inference Process”, In the Proceedings of 5th International Conference on Research and Development in Information Retrieval, 1993.
[ 3 ] Porter, M. “An Algorithm for Suffix Stripping”, Program, 14(3): 130-137, July 1980.
[ 4 ] Paik, J.H. and Parui, S.K. “A Simple Stemmer for Inflectional Languages,” Forum for Information Retrieval Evaluation, 2008.

No comments:

Post a Comment