Zero-shot object detection and referring expression comprehension using vision-language models
This project focused on constructing a comprehensive perception pipeline integrating Natural Language Processing (NLP), zero-shot object detection, and Referring Expression Comprehension (ReC) within a ROS (Robot Operating System) framework. The aim was to enhance robotic assistive devices in accura...
Saved in:
Main Author: | |
---|---|
Other Authors: | |
Format: | Final Year Project |
Language: | English |
Published: |
Nanyang Technological University
2024
|
Subjects: | |
Online Access: | https://hdl.handle.net/10356/177827 |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Institution: | Nanyang Technological University |
Language: | English |
id |
sg-ntu-dr.10356-177827 |
---|---|
record_format |
dspace |
spelling |
sg-ntu-dr.10356-1778272024-06-02T23:56:55Z Zero-shot object detection and referring expression comprehension using vision-language models A Manicka, Praveen Ang Wei Tech School of Mechanical and Aerospace Engineering Rehabilitation Research Institute of Singapore (RRIS) WTAng@ntu.edu.sg Computer and Information Science Engineering This project focused on constructing a comprehensive perception pipeline integrating Natural Language Processing (NLP), zero-shot object detection, and Referring Expression Comprehension (ReC) within a ROS (Robot Operating System) framework. The aim was to enhance robotic assistive devices in accurately interpreting natural language commands and grounding language to physical objects in the real world. To achieve this, we compared various combinations of zero-shot object detectors and ReC models, specifically specifically OWL-ViT and Grounding DINO for zero-shot object detection; and ReCLIP and GPT-4 for ReC. Our evaluation assessed the models' capabilities in counting, spatial reasoning, understanding superlatives, handling multiple instances, self-referential comprehension, and identifying household objects. The findings were showed that GPT-4 outperformed ReCLIP as for the purpose of ReC, and the combination of Grounding DINO and GPT-4 proved to be the best zero-shot object detector and ReC pair. Bachelor's degree 2024-05-31T12:13:12Z 2024-05-31T12:13:12Z 2024 Final Year Project (FYP) A Manicka, P. (2024). Zero-shot object detection and referring expression comprehension using vision-language models. Final Year Project (FYP), Nanyang Technological University, Singapore. https://hdl.handle.net/10356/177827 https://hdl.handle.net/10356/177827 en application/pdf Nanyang Technological University |
institution |
Nanyang Technological University |
building |
NTU Library |
continent |
Asia |
country |
Singapore Singapore |
content_provider |
NTU Library |
collection |
DR-NTU |
language |
English |
topic |
Computer and Information Science Engineering |
spellingShingle |
Computer and Information Science Engineering A Manicka, Praveen Zero-shot object detection and referring expression comprehension using vision-language models |
description |
This project focused on constructing a comprehensive perception pipeline integrating Natural Language Processing (NLP), zero-shot object detection, and Referring Expression Comprehension (ReC) within a ROS (Robot Operating System) framework. The aim was to enhance robotic assistive devices in accurately interpreting natural language commands and grounding language to physical objects in the real world. To achieve this, we compared various combinations of zero-shot object detectors and ReC models, specifically specifically OWL-ViT and Grounding DINO for zero-shot object detection; and ReCLIP and GPT-4 for ReC. Our evaluation assessed the models' capabilities in counting, spatial reasoning, understanding superlatives, handling multiple instances, self-referential comprehension, and identifying household objects. The findings were showed that GPT-4 outperformed ReCLIP as for the purpose of ReC, and the combination of Grounding DINO and GPT-4 proved to be the best zero-shot object detector and ReC pair. |
author2 |
Ang Wei Tech |
author_facet |
Ang Wei Tech A Manicka, Praveen |
format |
Final Year Project |
author |
A Manicka, Praveen |
author_sort |
A Manicka, Praveen |
title |
Zero-shot object detection and referring expression comprehension using vision-language models |
title_short |
Zero-shot object detection and referring expression comprehension using vision-language models |
title_full |
Zero-shot object detection and referring expression comprehension using vision-language models |
title_fullStr |
Zero-shot object detection and referring expression comprehension using vision-language models |
title_full_unstemmed |
Zero-shot object detection and referring expression comprehension using vision-language models |
title_sort |
zero-shot object detection and referring expression comprehension using vision-language models |
publisher |
Nanyang Technological University |
publishDate |
2024 |
url |
https://hdl.handle.net/10356/177827 |
_version_ |
1800916371569115136 |