AimigoTutor - tutoring application using multi-modal capabilities
Video captioning has been an up-and-coming research topic. Thanks to the recent advances in the performance of deep neural networks, especially with transformers, video captioning is seeing a huge potential improvement in accuracy and versatility. Most state-of-the-art video captioning models employ...
Saved in:
Main Author: | Nguyen, Viet Hoang |
---|---|
Other Authors: | Hanwang Zhang |
Format: | Final Year Project |
Language: | English |
Published: |
Nanyang Technological University
2024
|
Subjects: | |
Online Access: | https://hdl.handle.net/10356/175732 |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Institution: | Nanyang Technological University |
Language: | English |
Similar Items
-
Online multi-face tracking with multi-modality cascaded matching
by: Weng, Zhenyu, et al.
Published: (2024) -
Q-instruct: improving low-level visual abilities for multi-modality foundation models
by: Wu, Haoning, et al.
Published: (2024) -
FindMyTutor: An Android application for matching students and private tutors
by: Warit Taveekarn, et al.
Published: (2018) -
APPLICATIONS OF MULTI-MODAL ATOMIC FORCE MICROSCOPE IN 2D TRANSITION METAL DICHALCOGENIDES
by: WANG XINYUN
Published: (2021) -
Unified information fusion network for multi-modal RGB-D and RGB-T salient object detection
by: Gao, Wei, et al.
Published: (2021)