Error analysis of Chinese text segmentation using statistical approach

C.C. Yang; K.W. Li

Back

Conference proceeding

Error analysis of Chinese text segmentation using statistical approach

C.C. Yang and K.W. Li

Proceedings of the 2004 Joint ACM/IEEE Conference on Digital Libraries, 2004

2004

Additional Links

Abstract

Content based retrieval

Dictionaries

Error analysis

Indexing

Information retrieval

Mutual information

Natural languages

Research and development management

Software libraries

Systems engineering and theory

The Chinese text segmentation is important for the indexing of Chinese documents, which has significant impact on the performance of Chinese information retrieval. The statistical approach overcomes the limitations of the dictionary based approach. The statistical approach is developed by utilizing the statistical information about the association of adjacent characters in Chinese text collected from the Chinese corpus. Both known words and unknown words can be segmented by the statistical approach. However, errors may occur due to the limitation of the corpus. In this work, we have conducted the error analysis of two Chinese text segmentation techniques using statistical approach, namely, boundary detection and heuristic method. Such error analysis is useful for the future development of the automatic text segmentation of Chinese text or other text in oriental languages. It is also helpful to understand the impact of these errors on the information retrieval system in digital libraries.

Metrics

8 Record Views

Details

Title: Error analysis of Chinese text segmentation using statistical approach
Creators: C.C. Yang - Chinese University of Hong Kong
K.W. Li - Chinese University of Hong Kong
Publication Details: Proceedings of the 2004 Joint ACM/IEEE Conference on Digital Libraries, 2004
Conference: 2004 Joint ACM/IEEE Conference on Digital Libraries, 2004
Publisher: IEEE
Resource Type: Conference proceeding
Language: English
Academic Unit: Information Science (Informatics)
Identifiers: 991021855282604721

Error analysis of Chinese text segmentation using statistical approach

Additional Links

Abstract

Metrics

Details

Drexel University Social media