🚀Excited to release our 112-page study on math reasoning in visual contexts via #MathVista. For the first time, we provide both quantitative and qualitative evaluations of #GPT4V, #Bard, & 10 other models. 📄✨Full paper: lupantech.github.io/papers/arxiv23… 🔗Proj: mathvista.github.io 🔍 Key Insights: 1️⃣ #GPT4V achieves a 49.9% accuracy, notably surpassing #Bard by 15.1%. However, it's still 10.4% behind human performance. 2️⃣ #GPT4V exhibits an emergent ability of self-verification, enabling it to autonomously check and refine its outcomes in a single inference – a feature absent in other models. 3️⃣ #GPT4V highlights its potential through self-consistency and multi-turn human-AI dialogues. 📜 Arxiv: arxiv.org/abs/2310.02255 (updated soon) 🛠️ Code: github.com/lupantech/Math… 📊
@huggingfaceData: huggingface.co/datasets/AI4Ma… 🔍 Visualization: mathvista.github.io/#visualization 🏆 Leaderboard: mathvista.github.io/#leaderboard A massive shoutout to our outstanding team from
@uclanlp,
@uwnlp, and
@MSFTResearch:
@hbXNov, Tony Xia,
@liujc1998,
@ChunyuanLi,
@HannaHajishirzi,
@kelvinih,
@kaiwei_chang, Michel Galley, and
@JianfengGao0217🧵1/N