Error anaylsis of Chinese text segmentation using statistical approach

Christopher C. Yang; Kar Wing Li

doi:10.1145/996350.996410

Back

Conference proceeding

Error anaylsis of Chinese text segmentation using statistical approach

Christopher C. Yang and Kar Wing Li

Proceedings of the 4th ACM/IEEE-CS joint conference on Digital libraries

07 Jun 2004

DOI: https://doi.org/10.1145/996350.996410

Featured in Collection : UN Sustainable Development Goals @ Drexel

Additional Links

Abstract

Applied computing -- Computers in other domains -- Digital libraries and archives

Information systems -- Information retrieval -- Document representation

Information systems -- Information systems applications -- Digital libraries and archives

The Chinese text segmentation is important for the indexing of Chinese documents, which has significant impact on the performance of Chinese information retrieval. The statistical approach overcomes the limitations of the dictionary based approach. The statistical approach is developed by utilizing the statistical information about the association of adjacent characters in Chinese text collected from the Chinese corpus Both known words and unknown words can be segmented by the statistical approach. However, errors may occur due to the limitation of the corpus. In this work, we have conducted the error analysis of two Chinese text segmentation techniques using statistical approach, namely, boundary detection and heuristic method Such error analysis is useful for the future development of the automatic text segmentation of Chinese text or other text in oriental languages. It is also helpful to understand the impact of these errors on the information retrieval system in digital libraries.

Metrics

9 Record Views

3 citations in Web of Science

Details

Title: Error anaylsis of Chinese text segmentation using statistical approach
Creators: Christopher C. Yang - Chinese University of Hong Kong
Kar Wing Li - Chinese University of Hong Kong
Publication Details: Proceedings of the 4th ACM/IEEE-CS joint conference on Digital libraries
Conference: JCDL04: ACM/IEEE Joint Conference on Digital Libraries 2004 (2004)
Series: ACM Conferences
Publisher: ACM
Resource Type: Conference proceeding
Language: English
Academic Unit: Information Science
Web of Science ID: WOS:000222881400046
Other Identifier: 991021855280504721

UN Sustainable Development Goals (SDGs)

This publication has contributed to the advancement of the following goals:

Source: SDGs in the Output

InCites Highlights

Data related to this publication, from InCites Benchmarking & Analytics tool:

Web of Science research areas: Computer Science, Information Systems; Computer Science, Interdisciplinary Applications; Information Science & Library Science

Error anaylsis of Chinese text segmentation using statistical approach

Additional Links

Abstract

Metrics

Details

UN Sustainable Development Goals (SDGs)

InCites Highlights

Drexel University Social media