X (formerly Twitter)

🔥Thrilled to share our #NeurIPS2024 paper, “JourneyBench⚖️: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images” A follow-up to our previous work, HaloQuest 😇(x.com/lmthang/status…) JourneyBench is a comprehensive human-annotated benchmark of Imaginary (unusual or fictional) prompt-based generated images. It is designed to assess the model’s fine-grained multimodal reasoning abilities across five tasks: complementary multimodal chain of thought, multi-image VQA, imaginary image captioning, VQA with hallucination triggers, and fine-grained retrieval with sample-specific distractors. It rigorously tests multimodal understanding of current VLMs by presenting scenarios that are uncommon in the real world. This requires models to comprehend the visual information rigorously rather than relying on past biases that the model might have learned. 📽️Project: journeybench.github.io 📜ArXiv: arxiv.org/pdf/2409.12953 🧑‍💻Code: github.com/JourneyBench/J… #⃣Data: huggingface.co/JourneyBench GPT-4o’s performance accuracy on three of the tasks are: 62.18% on [Arithmetic Reasoning] Complementary Multimodal Chain-Of-Thought 57.89% on [Multi-image Reasoning] Multi-Image VQA 68.10% on [Visual Reasoning] VQA with Hallucination Triggers I am honored to deliver this work with my wonderful collaborators. Shih-fu Chang

@CUSEAS@kaiwei_chang

Chris Thomas

@VT_CS@XyouH

@hammad214

@RuiSun94013021

Read the original on x.com ↗