Evaluating vision-language models long-chain reasoning ability with multiple ground truths
With the recent advancements in vision-language models, many researchers start to evaluate their various zero-shot capabilities to answer questions given a video input. However, there has not been a standardised and “best practice” method to evaluate the quality of a model’s open-ended answer given...
Saved in:
Main Author: | Setiadharma, Christopher Arif |
---|---|
Other Authors: | Liu Ziwei |
Format: | Final Year Project |
Language: | English |
Published: |
Nanyang Technological University
2024
|
Subjects: | |
Online Access: | https://hdl.handle.net/10356/175186 |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Institution: | Nanyang Technological University |
Language: | English |
Similar Items
-
Learning language to symbol and language to vision mapping for visual grounding
by: He, Su, et al.
Published: (2022) -
Towards Ground Truthing Observations in Gray-Box Anomaly Detection
by: MING, Jiang, et al.
Published: (2011) -
ROME: Evaluating pre-trained vision-language models on reasoning beyond visual common sense
by: ZHOU, Kankan, et al.
Published: (2023) -
Learning to compose and reason with language tree structures for visual grounding
by: Hong, Richang, et al.
Published: (2022) -
Is the ground truth really accurate? Dataset purification for automated program repair
by: YANG, Deheng, et al.
Published: (2021)