Keyword Extraction from a Single Document using Word Co-occurrence Statistical Information

Authors

Yutaka Matsuo

Mitsuru Ishizuka

Published:

May 2003

Proceedings:

Proceedings of the Sixteenth International Florida Artificial Intelligence Research Society Conference (FLAIRS 2003)

Volume

Issue:

Proceedings of the Sixteenth International Florida Artificial Intelligence Research Society Conference (FLAIRS 2003)

Track:

All Papers

Downloads:

Download PDF

Abstract:

We present a new keyword extraction algorithm that applies to a single document without using a corpus. Frequent terms are extracted first, then a set of cooccurrence between each term and the frequent terms, i.e., occurrences in the same sentences, is generated. Co-occurrence distribution shows importance of a term in the document as follows. If probability distribution of co-occurrence between term a and the frequent terms is biased to a particular subset of frequent terms, then term a is likely to be a keyword. The degree of biases of distribution is measured by the χ2-measure. Our algorithm shows comparable performance to tfidf without using a corpus.

FLAIRS

Proceedings of the Sixteenth International Florida Artificial Intelligence Research Society Conference (FLAIRS 2003)

ISBN 978-1-57735-177-1

Published by The AAAI Press, Menlo Park, California.

Cookie	Duration	Description
cookielawinfo-checkbox-analytics	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Analytics".
cookielawinfo-checkbox-functional	11 months	The cookie is set by GDPR cookie consent to record the user consent for the cookies in the category "Functional".
cookielawinfo-checkbox-necessary	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookies is used to store the user consent for the cookies in the category "Necessary".
cookielawinfo-checkbox-others	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Other.
cookielawinfo-checkbox-performance	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Performance".
viewed_cookie_policy	11 months	The cookie is set by the GDPR Cookie Consent plugin and is used to store whether or not user has consented to the use of cookies. It does not store any personal data.