Extracting time expressions and named entities with constituent-based tagging schemes

Time expressions and named entities play important roles in data mining, information retrieval, and natural language processing. However, the conventional position-based tagging schemes (e.g., the BIO and BILOU schemes) that previous research used to model time expressions and named entities suffer...

Full description

Saved in:
Bibliographic Details
Main Authors: Zhong, Xiaoshi, Cambria, Erik, Hussain Amir
Other Authors: School of Computer Science and Engineering
Format: Article
Language:English
Published: 2022
Subjects:
Online Access:https://hdl.handle.net/10356/155193
Tags: Add Tag
No Tags, Be the first to tag this record!
Institution: Nanyang Technological University
Language: English
id sg-ntu-dr.10356-155193
record_format dspace
spelling sg-ntu-dr.10356-1551932022-02-17T08:09:48Z Extracting time expressions and named entities with constituent-based tagging schemes Zhong, Xiaoshi Cambria, Erik Hussain Amir School of Computer Science and Engineering Engineering::Computer science and engineering Inconsistent Tag Assignment Position-Based Tagging Scheme Time expressions and named entities play important roles in data mining, information retrieval, and natural language processing. However, the conventional position-based tagging schemes (e.g., the BIO and BILOU schemes) that previous research used to model time expressions and named entities suffer from the problem of inconsistent tag assignment. To overcome the problem of inconsistent tag assignment, we designed a new type of tagging schemes to model time expressions and named entities based on their constituents. Specifically, to model time expressions, we defined a constituent-based tagging scheme termed TOMN scheme with four tags, namely T, O, M, and N, indicating the defined constituents of time expressions, namely time token, modifier, numeral, and the words outside time expressions. To model named entities, we defined a constituent-based tagging scheme termed UGTO scheme with four tags, namely U, G, T, and O, indicating the defined constituents of named entities, namely uncommon word, general modifier, trigger word, and the words outside named entities. In modeling, our TOMN and UGTO schemes model time expressions and named entities under conditional random fields with minimal features according to an in-depth analysis for the characteristics of time expressions and named entities. Experiments on diverse datasets demonstrate that our proposed methods perform equally with or more effectively than representative state-of-the-art methods on both time expression extraction and named entity extraction. Agency for Science, Technology and Research (A*STAR) This research was funded by the AME Programmatic Funding (Project No. A18A2b0046) from the Agency for Science, Technology and Research (A*STAR), Singapore. 2022-02-16T07:43:34Z 2022-02-16T07:43:34Z 2020 Journal Article Zhong, X., Cambria, E. & Hussain Amir (2020). Extracting time expressions and named entities with constituent-based tagging schemes. Cognitive Computation, 12(4), 844-862. https://dx.doi.org/10.1007/s12559-020-09714-8 1866-9956 https://hdl.handle.net/10356/155193 10.1007/s12559-020-09714-8 2-s2.0-85084445501 4 12 844 862 en A18A2b0046 Cognitive Computation © 2020 Springer Science+Business Media, LLC, part of Springer Nature. All rights reserved.
institution Nanyang Technological University
building NTU Library
continent Asia
country Singapore
Singapore
content_provider NTU Library
collection DR-NTU
language English
topic Engineering::Computer science and engineering
Inconsistent Tag Assignment
Position-Based Tagging Scheme
spellingShingle Engineering::Computer science and engineering
Inconsistent Tag Assignment
Position-Based Tagging Scheme
Zhong, Xiaoshi
Cambria, Erik
Hussain Amir
Extracting time expressions and named entities with constituent-based tagging schemes
description Time expressions and named entities play important roles in data mining, information retrieval, and natural language processing. However, the conventional position-based tagging schemes (e.g., the BIO and BILOU schemes) that previous research used to model time expressions and named entities suffer from the problem of inconsistent tag assignment. To overcome the problem of inconsistent tag assignment, we designed a new type of tagging schemes to model time expressions and named entities based on their constituents. Specifically, to model time expressions, we defined a constituent-based tagging scheme termed TOMN scheme with four tags, namely T, O, M, and N, indicating the defined constituents of time expressions, namely time token, modifier, numeral, and the words outside time expressions. To model named entities, we defined a constituent-based tagging scheme termed UGTO scheme with four tags, namely U, G, T, and O, indicating the defined constituents of named entities, namely uncommon word, general modifier, trigger word, and the words outside named entities. In modeling, our TOMN and UGTO schemes model time expressions and named entities under conditional random fields with minimal features according to an in-depth analysis for the characteristics of time expressions and named entities. Experiments on diverse datasets demonstrate that our proposed methods perform equally with or more effectively than representative state-of-the-art methods on both time expression extraction and named entity extraction.
author2 School of Computer Science and Engineering
author_facet School of Computer Science and Engineering
Zhong, Xiaoshi
Cambria, Erik
Hussain Amir
format Article
author Zhong, Xiaoshi
Cambria, Erik
Hussain Amir
author_sort Zhong, Xiaoshi
title Extracting time expressions and named entities with constituent-based tagging schemes
title_short Extracting time expressions and named entities with constituent-based tagging schemes
title_full Extracting time expressions and named entities with constituent-based tagging schemes
title_fullStr Extracting time expressions and named entities with constituent-based tagging schemes
title_full_unstemmed Extracting time expressions and named entities with constituent-based tagging schemes
title_sort extracting time expressions and named entities with constituent-based tagging schemes
publishDate 2022
url https://hdl.handle.net/10356/155193
_version_ 1725985746855133184