Chinese words segmentation in user generated content
Chinese word segmentation is the first step for Chinese text processing. The accuracy of Chinese word segmentation directly affects the performance of Chinese text processing. Therefore, Chinese word segmentation plays an important role in Chinese text processing. In addition, with the increasing po...
Saved in:
Main Author: | |
---|---|
Other Authors: | |
Format: | Final Year Project |
Language: | English |
Published: |
2015
|
Subjects: | |
Online Access: | http://hdl.handle.net/10356/63803 |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Institution: | Nanyang Technological University |
Language: | English |
id |
sg-ntu-dr.10356-63803 |
---|---|
record_format |
dspace |
spelling |
sg-ntu-dr.10356-638032023-03-03T20:24:01Z Chinese words segmentation in user generated content Cai, Xiaoxuan Sun Aixin School of Computer Engineering DRNTU::Engineering::Computer science and engineering Chinese word segmentation is the first step for Chinese text processing. The accuracy of Chinese word segmentation directly affects the performance of Chinese text processing. Therefore, Chinese word segmentation plays an important role in Chinese text processing. In addition, with the increasing popularity of social media in China, Chinese sentences that are written in an informal manner in user generated content are very common on the Internet. This project is to study Chinese word segmentation in user generated content. In this project, two existing Chinese word segmentation tools Jieba [1] and Stanford Word Segmenter [2] are studied; a new Chinese word segmentation tool named Weibo Segmenter implemented according to [3] is presented; then these three tools are tested using the same dataset to compare the performance. As a result, Weibo Segmenter achieves an accuracy rate of 83.3% in the test. The performance of Weibo Segmenter could be further enhanced by using a more suitable dictionary and some programming techniques. Bachelor of Engineering (Computer Science) 2015-05-19T03:54:24Z 2015-05-19T03:54:24Z 2015 2015 Final Year Project (FYP) http://hdl.handle.net/10356/63803 en Nanyang Technological University 42 p. application/pdf |
institution |
Nanyang Technological University |
building |
NTU Library |
continent |
Asia |
country |
Singapore Singapore |
content_provider |
NTU Library |
collection |
DR-NTU |
language |
English |
topic |
DRNTU::Engineering::Computer science and engineering |
spellingShingle |
DRNTU::Engineering::Computer science and engineering Cai, Xiaoxuan Chinese words segmentation in user generated content |
description |
Chinese word segmentation is the first step for Chinese text processing. The accuracy of Chinese word segmentation directly affects the performance of Chinese text processing. Therefore, Chinese word segmentation plays an important role in Chinese text processing. In addition, with the increasing popularity of social media in China, Chinese sentences that are written in an informal manner in user generated content are very common on the Internet. This project is to study Chinese word segmentation in user generated content. In this project, two existing Chinese word segmentation tools Jieba [1] and Stanford Word Segmenter [2] are studied; a new Chinese word segmentation tool named Weibo Segmenter implemented according to [3] is presented; then these three tools are tested using the same dataset to compare the performance. As a result, Weibo Segmenter achieves an accuracy rate of 83.3% in the test. The performance of Weibo Segmenter could be further enhanced by using a more suitable dictionary and some programming techniques. |
author2 |
Sun Aixin |
author_facet |
Sun Aixin Cai, Xiaoxuan |
format |
Final Year Project |
author |
Cai, Xiaoxuan |
author_sort |
Cai, Xiaoxuan |
title |
Chinese words segmentation in user generated content |
title_short |
Chinese words segmentation in user generated content |
title_full |
Chinese words segmentation in user generated content |
title_fullStr |
Chinese words segmentation in user generated content |
title_full_unstemmed |
Chinese words segmentation in user generated content |
title_sort |
chinese words segmentation in user generated content |
publishDate |
2015 |
url |
http://hdl.handle.net/10356/63803 |
_version_ |
1759854911201214464 |