Emergent semantic segmentation: training-free dense-label-free extraction from vision-language models

From an enormous amount of image-text pairs, large-scale vision-language models (VLMs) learn to implicitly associate image regions with words, which is vital for tasks such as image captioning and visual question answering. However, leveraging such pre-trained models for open-vocabulary semantic s...

Full description

Saved in:

Bibliographic Details
Main Author:	Luo, Jiayun
Other Authors:	Li Boyang
Format:	Thesis-Master by Research
Language:	English
Published:	Nanyang Technological University 2024
Subjects:	Computer and Information Science Vision-language model Open-vocabulary semantic segementation
Online Access:	https://hdl.handle.net/10356/175765
Tags:	Add Tag No Tags, Be the first to tag this record!
Institution:	Nanyang Technological University
Language:	English

Be the first to leave a comment!

Emergent semantic segmentation: training-free dense-label-free extraction from vision-language models

Similar Items