vqa论文

Searching…

arxiv.org

[2501.03939] Visual question answering: from early ...

Jan 7, 2025 · We review major VQA approaches, focusing on deep learning-based methods, and explore the emerging field of Large Visual Language Models (LVLMs) that have demonstrated success in multimodal tasks like VQA.

blog.csdn.net

LLM | 论文精读 | CVPR | 基于问题驱动图像描述的视觉问答增强引言_en...

Nov 8, 2024 · 本文提出了一种增强视觉问答(VQA)性能的新方法,通过生成问题驱动的图像描述作为中间步骤,将上下文信息有效融入到问答过程中,尤其在零样本场景中展现出显著的优势。 研究通过关键词提取技术使描述与问题紧密结合,从而提高了模型的理解和推理能力。