X (formerly Twitter)

🚀Excited to release our 112-page study on math reasoning in visual contexts via #MathVista. For the first time, we provide both quantitative and qualitative evaluations of #GPT4V, #Bard, & 10 other models. 📄✨Full paper: lupantech.github.io/papers/arxiv23… 🔗Proj: mathvista.github.io 🔍 Key Insights: 1️⃣ #GPT4V achieves a 49.9% accuracy, notably surpassing #Bard by 15.1%. However, it's still 10.4% behind human performance. 2️⃣ #GPT4V exhibits an emergent ability of self-verification, enabling it to autonomously check and refine its outcomes in a single inference – a feature absent in other models. 3️⃣ #GPT4V highlights its potential through self-consistency and multi-turn human-AI dialogues. 📜 Arxiv: arxiv.org/abs/2310.02255 (updated soon) 🛠️ Code: github.com/lupantech/Math… 📊

@huggingface

Data: huggingface.co/datasets/AI4Ma… 🔍 Visualization: mathvista.github.io/#visualization 🏆 Leaderboard: mathvista.github.io/#leaderboard A massive shoutout to our outstanding team from

@uclanlp

,

@uwnlp

, and

@MSFTResearch

:

@hbXNov

, Tony Xia,

@liujc1998

,

@ChunyuanLi

,

@HannaHajishirzi

,

@kelvinih

,

@kaiwei_chang

, Michel Galley, and

@JianfengGao0217

🧵1/N

Read the original on x.com ↗