Emergent semantic segmentation: training-free dense-label-free extraction from vision-language models

From an enormous amount of image-text pairs, large-scale vision-language models (VLMs) learn to implicitly associate image regions with words, which is vital for tasks such as image captioning and visual question answering. However, leveraging such pre-trained models for open-vocabulary semantic s...

全面介紹

Saved in:

書目詳細資料
主要作者:	Luo, Jiayun
其他作者:	Li Boyang
格式:	Thesis-Master by Research
語言:	English
出版:	Nanyang Technological University 2024
主題:	Computer and Information Science Vision-language model Open-vocabulary semantic segementation
在線閱讀:	https://hdl.handle.net/10356/175765
標簽:	添加標簽沒有標簽, 成為第一個標記此記錄!
機構:	Nanyang Technological University
語言:	English

因特網

https://hdl.handle.net/10356/175765

Emergent semantic segmentation: training-free dense-label-free extraction from vision-language models

因特網

相似書籍