Building the foundation text for Nanyang Technological University : multilingual corpus (NTU-MC)

The NTU-MC is a multilingual corpus that taps on the availability of multilingual text available in Singapore. The current version of NTU-MC contains a total of ~375,000 words (15,096 sentences) for the NTU-MC in 6 languages (English, Chinese, Japanese, Korean, Indonesian and Vietnamese) from 6 lang...

Full description

Saved in:
Bibliographic Details
Main Author: Tan, Li Ling.
Other Authors: Francis Bond
Format: Final Year Project
Language:English
Published: 2012
Subjects:
Online Access:https://hdl.handle.net/10356/92348
http://hdl.handle.net/10220/7790
Tags: Add Tag
No Tags, Be the first to tag this record!
Institution: Nanyang Technological University
Language: English
id sg-ntu-dr.10356-92348
record_format dspace
spelling sg-ntu-dr.10356-923482021-12-20T01:12:58Z Building the foundation text for Nanyang Technological University : multilingual corpus (NTU-MC) Tan, Li Ling. Francis Bond School of Humanities and Social Sciences DRNTU::Humanities::Language::Linguistics The NTU-MC is a multilingual corpus that taps on the availability of multilingual text available in Singapore. The current version of NTU-MC contains a total of ~375,000 words (15,096 sentences) for the NTU-MC in 6 languages (English, Chinese, Japanese, Korean, Indonesian and Vietnamese) from 6 language families (Indo-European, Japonic, Austro-Asiatic, Sino-Tibetan, Austronesian and Korean as a language isolate); all text in English, Chinese, Japanese, Korean and Vietnamese were Part Of Speech (POS) tagged. This project focuses on compiling the foundation text for the NTU-MC and this dissertation describes the motivations, the corpus compilation process and internal and cross-corpora evaluation of the corpus output. The corpus will be made available to the public under the Creative Common – Attribute 3.0 Unported license in Summer 2011. Bachelor of Arts in Linguistics and Multilingual Studies 2012-04-13T04:46:34Z 2019-12-06T18:21:45Z 2012-04-13T04:46:34Z 2019-12-06T18:21:45Z 2011 2011 Final Year Project (FYP) Tan, L. L. (2011). Building the Foundation Text for Nanyang Technological University : Multilingual Corpus (NTU-MC). Final year project report, Nanyang Technological University. https://hdl.handle.net/10356/92348 http://hdl.handle.net/10220/7790 en 47 p. application/pdf
institution Nanyang Technological University
building NTU Library
continent Asia
country Singapore
Singapore
content_provider NTU Library
collection DR-NTU
language English
topic DRNTU::Humanities::Language::Linguistics
spellingShingle DRNTU::Humanities::Language::Linguistics
Tan, Li Ling.
Building the foundation text for Nanyang Technological University : multilingual corpus (NTU-MC)
description The NTU-MC is a multilingual corpus that taps on the availability of multilingual text available in Singapore. The current version of NTU-MC contains a total of ~375,000 words (15,096 sentences) for the NTU-MC in 6 languages (English, Chinese, Japanese, Korean, Indonesian and Vietnamese) from 6 language families (Indo-European, Japonic, Austro-Asiatic, Sino-Tibetan, Austronesian and Korean as a language isolate); all text in English, Chinese, Japanese, Korean and Vietnamese were Part Of Speech (POS) tagged. This project focuses on compiling the foundation text for the NTU-MC and this dissertation describes the motivations, the corpus compilation process and internal and cross-corpora evaluation of the corpus output. The corpus will be made available to the public under the Creative Common – Attribute 3.0 Unported license in Summer 2011.
author2 Francis Bond
author_facet Francis Bond
Tan, Li Ling.
format Final Year Project
author Tan, Li Ling.
author_sort Tan, Li Ling.
title Building the foundation text for Nanyang Technological University : multilingual corpus (NTU-MC)
title_short Building the foundation text for Nanyang Technological University : multilingual corpus (NTU-MC)
title_full Building the foundation text for Nanyang Technological University : multilingual corpus (NTU-MC)
title_fullStr Building the foundation text for Nanyang Technological University : multilingual corpus (NTU-MC)
title_full_unstemmed Building the foundation text for Nanyang Technological University : multilingual corpus (NTU-MC)
title_sort building the foundation text for nanyang technological university : multilingual corpus (ntu-mc)
publishDate 2012
url https://hdl.handle.net/10356/92348
http://hdl.handle.net/10220/7790
_version_ 1720447109620039680