Tree-based text stream clustering with application to spam mail classification

Copyright © 2018 Inderscience Enterprises Ltd. This paper proposes a new text clustering algorithm based on a tree structure. The main idea of the clustering algorithm is a sub-tree at a specific node represents a document cluster. Our clustering algorithm is a single pass scanning algorithm which t...

Full description

Saved in:
Bibliographic Details
Main Authors: Phimphaka Taninpong, Sudsanguan Ngamsuriyaroj
Other Authors: Mahidol University
Format: Article
Published: 2019
Subjects:
Online Access:https://repository.li.mahidol.ac.th/handle/123456789/45390
Tags: Add Tag
No Tags, Be the first to tag this record!
Institution: Mahidol University
id th-mahidol.45390
record_format dspace
spelling th-mahidol.453902019-08-23T18:31:11Z Tree-based text stream clustering with application to spam mail classification Phimphaka Taninpong Sudsanguan Ngamsuriyaroj Mahidol University Business, Management and Accounting Computer Science Mathematics Copyright © 2018 Inderscience Enterprises Ltd. This paper proposes a new text clustering algorithm based on a tree structure. The main idea of the clustering algorithm is a sub-tree at a specific node represents a document cluster. Our clustering algorithm is a single pass scanning algorithm which traverses down the tree to search for all clusters without having to predefine the number of clusters. Thus, it fits our objectives to produce document clusters having high cohesion, and to keep the minimum number of clusters. Moreover, an incremental learning process will perform after a new document is inserted into the tree, and the clusters will be rebuilt to accommodate the new information. In addition, we applied the proposed clustering algorithm to spam mail classification and the experimental results show that tree-based text clustering spam filter gives higher accuracy and specificity than the cobweb clustering, naïve Bayes and KNN. 2019-08-23T10:43:37Z 2019-08-23T10:43:37Z 2018-01-01 Article International Journal of Data Mining, Modelling and Management. Vol.10, No.4 (2018), 353-370 10.1504/IJDMMM.2018.095354 17591171 17591163 2-s2.0-85054534251 https://repository.li.mahidol.ac.th/handle/123456789/45390 Mahidol University SCOPUS https://www.scopus.com/inward/record.uri?partnerID=HzOxMe3b&scp=85054534251&origin=inward
institution Mahidol University
building Mahidol University Library
continent Asia
country Thailand
Thailand
content_provider Mahidol University Library
collection Mahidol University Institutional Repository
topic Business, Management and Accounting
Computer Science
Mathematics
spellingShingle Business, Management and Accounting
Computer Science
Mathematics
Phimphaka Taninpong
Sudsanguan Ngamsuriyaroj
Tree-based text stream clustering with application to spam mail classification
description Copyright © 2018 Inderscience Enterprises Ltd. This paper proposes a new text clustering algorithm based on a tree structure. The main idea of the clustering algorithm is a sub-tree at a specific node represents a document cluster. Our clustering algorithm is a single pass scanning algorithm which traverses down the tree to search for all clusters without having to predefine the number of clusters. Thus, it fits our objectives to produce document clusters having high cohesion, and to keep the minimum number of clusters. Moreover, an incremental learning process will perform after a new document is inserted into the tree, and the clusters will be rebuilt to accommodate the new information. In addition, we applied the proposed clustering algorithm to spam mail classification and the experimental results show that tree-based text clustering spam filter gives higher accuracy and specificity than the cobweb clustering, naïve Bayes and KNN.
author2 Mahidol University
author_facet Mahidol University
Phimphaka Taninpong
Sudsanguan Ngamsuriyaroj
format Article
author Phimphaka Taninpong
Sudsanguan Ngamsuriyaroj
author_sort Phimphaka Taninpong
title Tree-based text stream clustering with application to spam mail classification
title_short Tree-based text stream clustering with application to spam mail classification
title_full Tree-based text stream clustering with application to spam mail classification
title_fullStr Tree-based text stream clustering with application to spam mail classification
title_full_unstemmed Tree-based text stream clustering with application to spam mail classification
title_sort tree-based text stream clustering with application to spam mail classification
publishDate 2019
url https://repository.li.mahidol.ac.th/handle/123456789/45390
_version_ 1763488365069467648