Maxenttagger for Malays Jawi POS-tags
Purpose - Malay is a major language of the Austronesian family spoken in many countries. Malay Jawi is lacking in annotated resources and tools. In addition, Part-of-speech (POS) ambiguity in Natural Language Processing (NLP) is a vague important phenomenon that needs to be solved immediately. Sinc...
Saved in:
Main Authors: | , , , |
---|---|
Format: | Conference or Workshop Item |
Language: | English |
Published: |
2017
|
Subjects: | |
Online Access: | http://repo.uum.edu.my/24498/1/SICONSEM%202017%2031%2033.pdf http://repo.uum.edu.my/24498/ |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Institution: | Universiti Utara Malaysia |
Language: | English |
id |
my.uum.repo.24498 |
---|---|
record_format |
eprints |
spelling |
my.uum.repo.244982018-07-30T01:10:36Z http://repo.uum.edu.my/24498/ Maxenttagger for Malays Jawi POS-tags Abu Bakar, Juhaida Omar, Khairuddin Nasrudin, Mohammad Faidzul Murah, Mohd Zamri QA75 Electronic computers. Computer science Purpose - Malay is a major language of the Austronesian family spoken in many countries. Malay Jawi is lacking in annotated resources and tools. In addition, Part-of-speech (POS) ambiguity in Natural Language Processing (NLP) is a vague important phenomenon that needs to be solved immediately. Since POS is an important feature of the word, and is the link between the words and syntax, POS tagging (POST) needs to provide intermediate results showing superior performance to the next NLP tasks. POS ambiguity is a main problem in increasing POST performance. POST performance is often measured with accuracy and precision of a tag and it was considered critical to NLP application. Some of the standard package POS tagging provided in Natural Language ToolKit (NLTK) are Brill tagger, HMM tagger, and CRF Tagger. In this paper, POST Malay Jawi implemented NLP tools, NLTK for the state-of-the-art methods tagger; maximum entropy models. NLTK is used as the implementation tool for Jawi tagging, as syntax and semantics of the language is transparent, and it has the good functionality of NLP-operator. The tool also uses Python as the implementation language. 2017-12-04 Conference or Workshop Item PeerReviewed application/pdf en http://repo.uum.edu.my/24498/1/SICONSEM%202017%2031%2033.pdf Abu Bakar, Juhaida and Omar, Khairuddin and Nasrudin, Mohammad Faidzul and Murah, Mohd Zamri (2017) Maxenttagger for Malays Jawi POS-tags. In: Sintok International Conference on Social Science and Management (SICONSEM 2017), 4-5 December 2017, Adya Hotel, Langkawi Island, Kedah, Malaysia. |
institution |
Universiti Utara Malaysia |
building |
UUM Library |
collection |
Institutional Repository |
continent |
Asia |
country |
Malaysia |
content_provider |
Universiti Utara Malaysia |
content_source |
UUM Institutionali Repository |
url_provider |
http://repo.uum.edu.my/ |
language |
English |
topic |
QA75 Electronic computers. Computer science |
spellingShingle |
QA75 Electronic computers. Computer science Abu Bakar, Juhaida Omar, Khairuddin Nasrudin, Mohammad Faidzul Murah, Mohd Zamri Maxenttagger for Malays Jawi POS-tags |
description |
Purpose - Malay is a major language of the Austronesian family spoken in many countries. Malay Jawi is lacking in annotated resources and tools. In addition, Part-of-speech (POS) ambiguity in Natural Language Processing (NLP) is a vague important phenomenon that needs to be solved
immediately. Since POS is an important feature of the word, and is the link between the words and
syntax, POS tagging (POST) needs to provide intermediate results showing superior performance
to the next NLP tasks. POS ambiguity is a main problem in increasing POST performance. POST
performance is often measured with accuracy and precision of a tag and it was considered critical
to NLP application. Some of the standard package POS tagging provided in Natural Language
ToolKit (NLTK) are Brill tagger, HMM tagger, and CRF Tagger. In this paper, POST Malay Jawi
implemented NLP tools, NLTK for the state-of-the-art methods tagger; maximum entropy models.
NLTK is used as the implementation tool for Jawi tagging, as syntax and semantics of the language
is transparent, and it has the good functionality of NLP-operator. The tool also uses Python as the
implementation language. |
format |
Conference or Workshop Item |
author |
Abu Bakar, Juhaida Omar, Khairuddin Nasrudin, Mohammad Faidzul Murah, Mohd Zamri |
author_facet |
Abu Bakar, Juhaida Omar, Khairuddin Nasrudin, Mohammad Faidzul Murah, Mohd Zamri |
author_sort |
Abu Bakar, Juhaida |
title |
Maxenttagger for Malays Jawi POS-tags |
title_short |
Maxenttagger for Malays Jawi POS-tags |
title_full |
Maxenttagger for Malays Jawi POS-tags |
title_fullStr |
Maxenttagger for Malays Jawi POS-tags |
title_full_unstemmed |
Maxenttagger for Malays Jawi POS-tags |
title_sort |
maxenttagger for malays jawi pos-tags |
publishDate |
2017 |
url |
http://repo.uum.edu.my/24498/1/SICONSEM%202017%2031%2033.pdf http://repo.uum.edu.my/24498/ |
_version_ |
1644284067915497472 |