Automatic document summarization from social media and online news
This dissertation provides a new method for sentence embedding and document summarization. The topic model is utilized to modify the sentence embedding method SIF by capturing the information in the document, instead of relying on an external corpus. Thus, the modification embeds the information of...
Saved in:
主要作者: | |
---|---|
其他作者: | |
格式: | Thesis-Master by Coursework |
語言: | English |
出版: |
Nanyang Technological University
2020
|
主題: | |
在線閱讀: | https://hdl.handle.net/10356/141154 |
標簽: |
添加標簽
沒有標簽, 成為第一個標記此記錄!
|
機構: | Nanyang Technological University |
語言: | English |
總結: | This dissertation provides a new method for sentence embedding and document summarization. The topic model is utilized to modify the sentence embedding method SIF by capturing the information in the document, instead of relying on an external corpus. Thus, the modification embeds the information of the entire document into the sentence vectors, which is beneficial for further information extraction. Then we employ the graph-based method to score the sentences and select the high-scoring sentences to form the summary. In addition, this dissertation also tested the impact of different parameter changes in the model. The experimental results show that the proposed model can beat other classic and advanced models in semantic analysis and summary extraction with strong robustness. The datasets used in this dissertation are from social media and online news, which proves the applicability of this model to online information extraction. |
---|