Temporal sentence grounding in videos: a survey and future directions
Temporal sentence grounding in videos (TSGV), a.k.a., natural language video localization (NLVL) or video moment retrieval (VMR), aims to retrieve a temporal moment that semantically corresponds to a language query from an untrimmed video. Connecting computer vision and natural language, TSGV has dr...
Saved in:
Main Authors: | Zhang, Hao, Sun, Aixin, Jing, Wei, Zhou, Joey Tianyi |
---|---|
Other Authors: | School of Computer Science and Engineering |
Format: | Article |
Language: | English |
Published: |
2023
|
Subjects: | |
Online Access: | https://hdl.handle.net/10356/172187 |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Institution: | Nanyang Technological University |
Language: | English |
Similar Items
-
Efficient cross-modal video retrieval with meta-optimized frames
by: HAN, Ning, et al.
Published: (2024) -
Cross-modal Moment Localization in Videos
by: Meng Liu, et al.
Published: (2020) -
Towards temporal sentence grounding in videos
by: Zhang, Hao
Published: (2022) -
Learning a cross-modal hashing network for multimedia search
by: Tan, Yap Peng, et al.
Published: (2018) -
Deep Understanding of Cooking Procedure for Cross-modal Recipe Retrieval
by: Jing-Jing Chen, et al.
Published: (2020)