Training deep network models for accurate recognition of texts in scenes

Deep learning has seen a resurgence in the machine learning community in the past decade. Research on scene text detection and recognition using deep learning allows for more innovation in solving current issues. Current solutions treat text recognition as a category to be researched separately and...

Full description

Saved in:
Bibliographic Details
Main Author: Teo, Ren Jie
Other Authors: Lu Shijian
Format: Final Year Project
Language:English
Published: Nanyang Technological University 2020
Subjects:
Online Access:https://hdl.handle.net/10356/138112
Tags: Add Tag
No Tags, Be the first to tag this record!
Institution: Nanyang Technological University
Language: English
Description
Summary:Deep learning has seen a resurgence in the machine learning community in the past decade. Research on scene text detection and recognition using deep learning allows for more innovation in solving current issues. Current solutions treat text recognition as a category to be researched separately and more could be done to improve on that,making it an end-to-end system. In this FYP, deep learning network models will be implemented to recognise texts in various scenes. Hyper parameters of the deep learning network models will be fine-tuned to achieve optimal performance. This will create a great learning and practical experience. For optimal performance, it may be necessary to give up test accuracy for training speed sometimes. Comparison will be made between the deep learning network model fine-tuned for optimal performance, and a recent state-of-the-art deep learning network model without fine-tuning. This shows the improvement in research on the subject area. However,without a text detection system working in tandem with the text recognition model,scene text recognition will not serve much real world use. Future work for this project recommends better hardware to allow for more room to work with when fine-tuning hyper parameters, and possible integration with another system to make the scene text recognition an end-to-end model.