Automatic document summarization from social media and online news

This dissertation provides a new method for sentence embedding and document summarization. The topic model is utilized to modify the sentence embedding method SIF by capturing the information in the document, instead of relying on an external corpus. Thus, the modification embeds the information of...

Full description

Saved in:
Bibliographic Details
Main Author: Feng, Zijian
Other Authors: Mao Kezhi
Format: Thesis-Master by Coursework
Language:English
Published: Nanyang Technological University 2020
Subjects:
Online Access:https://hdl.handle.net/10356/141154
Tags: Add Tag
No Tags, Be the first to tag this record!
Institution: Nanyang Technological University
Language: English
Description
Summary:This dissertation provides a new method for sentence embedding and document summarization. The topic model is utilized to modify the sentence embedding method SIF by capturing the information in the document, instead of relying on an external corpus. Thus, the modification embeds the information of the entire document into the sentence vectors, which is beneficial for further information extraction. Then we employ the graph-based method to score the sentences and select the high-scoring sentences to form the summary. In addition, this dissertation also tested the impact of different parameter changes in the model. The experimental results show that the proposed model can beat other classic and advanced models in semantic analysis and summary extraction with strong robustness. The datasets used in this dissertation are from social media and online news, which proves the applicability of this model to online information extraction.