Algorithm of the longest commonly consecutive word for plagiarism detection in text based document
Plagiarism is a form of academic misconduct which has increased with the easy access to obtain information through electronic documents and the Internet. The problem of finding document plagiarism in full text document can be viewed as a problem of finding the longest common parts of strings.Moreove...
Saved in:
Main Authors: | , |
---|---|
Format: | Conference or Workshop Item |
Language: | English |
Published: |
2008
|
Subjects: | |
Online Access: | http://repo.uum.edu.my/9254/1/3.pdf http://repo.uum.edu.my/9254/ http://dx.doi.org/10.1109/ICDIM.2008.4746827 |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Institution: | Universiti Utara Malaysia |
Language: | English |
id |
my.uum.repo.9254 |
---|---|
record_format |
eprints |
spelling |
my.uum.repo.92542013-10-27T06:11:42Z http://repo.uum.edu.my/9254/ Algorithm of the longest commonly consecutive word for plagiarism detection in text based document Sediyono, Agung Ku-Mahamud, Ku Ruhana QA Mathematics Plagiarism is a form of academic misconduct which has increased with the easy access to obtain information through electronic documents and the Internet. The problem of finding document plagiarism in full text document can be viewed as a problem of finding the longest common parts of strings.Moreover, the detection system has to be capable to determine and visualize not only the common parts but also the location of the common parts in both the source and the observed document. Unlike previous research, this paper proposes a numerical based comparison algorithm that is comparable in the computation time without loosing the word order of common parts. Based on the experiment, the proposed algorithm outperforms the suffix tree in the length of observed paragraph below one hundred words. 2008 Conference or Workshop Item PeerReviewed application/pdf en http://repo.uum.edu.my/9254/1/3.pdf Sediyono, Agung and Ku-Mahamud, Ku Ruhana (2008) Algorithm of the longest commonly consecutive word for plagiarism detection in text based document. In: Third International Conference on Digital Information Management, 2008 ( ICDIM 2008), 13-16 Nov. 2008 , London. http://dx.doi.org/10.1109/ICDIM.2008.4746827 doi:10.1109/ICDIM.2008.4746827 |
institution |
Universiti Utara Malaysia |
building |
UUM Library |
collection |
Institutional Repository |
continent |
Asia |
country |
Malaysia |
content_provider |
Universiti Utara Malaysia |
content_source |
UUM Institutionali Repository |
url_provider |
http://repo.uum.edu.my/ |
language |
English |
topic |
QA Mathematics |
spellingShingle |
QA Mathematics Sediyono, Agung Ku-Mahamud, Ku Ruhana Algorithm of the longest commonly consecutive word for plagiarism detection in text based document |
description |
Plagiarism is a form of academic misconduct which has increased with the easy access to obtain information through electronic documents and the Internet. The problem of finding document plagiarism in full text document can be viewed as a problem of finding the longest common parts of strings.Moreover, the detection system has to be capable to determine and visualize not only the common parts but also the location of the common parts in both the source and the observed document. Unlike previous research, this paper proposes a numerical based comparison algorithm that is comparable in the computation time without loosing the word order of common parts. Based on the experiment, the proposed algorithm outperforms the suffix tree in the length of observed paragraph below one hundred words. |
format |
Conference or Workshop Item |
author |
Sediyono, Agung Ku-Mahamud, Ku Ruhana |
author_facet |
Sediyono, Agung Ku-Mahamud, Ku Ruhana |
author_sort |
Sediyono, Agung |
title |
Algorithm of the longest commonly consecutive word for plagiarism detection in text based document |
title_short |
Algorithm of the longest commonly consecutive word for plagiarism detection in text based document |
title_full |
Algorithm of the longest commonly consecutive word for plagiarism detection in text based document |
title_fullStr |
Algorithm of the longest commonly consecutive word for plagiarism detection in text based document |
title_full_unstemmed |
Algorithm of the longest commonly consecutive word for plagiarism detection in text based document |
title_sort |
algorithm of the longest commonly consecutive word for plagiarism detection in text based document |
publishDate |
2008 |
url |
http://repo.uum.edu.my/9254/1/3.pdf http://repo.uum.edu.my/9254/ http://dx.doi.org/10.1109/ICDIM.2008.4746827 |
_version_ |
1644280058671529984 |