RSS Amplifier

The Pennsylvania Heretic · Feb 5, 2025

A brief update on o3-mini and Starburst

0
Sign in to vote or save

Chapin Lenthall-Cleary · The Pennsylvania Heretic

I previously tested a variety of LLMs on Starburst, a game where one tries to figure out the laws of physics in a simple imaginary universe. All, including OpenAI's o1, performed poorly. o3-mini-high, the best currently available model from OpenAI, still performs poorly (at my estimate, very roughly a 2.7 intelligence on a 1-5 scale, compared to 2.6 for o1 and 2.3-2.5 for other contemporary models), though it demonstrates a modest improvement.

Though they're useful (and dangerous) in some ways and only becoming moreso, I remain skeptical about the abilities of LLMs to do true deep, understanding-based reasoning. On Starburst (and certain other benchmarks that can't be defeated by throwing dumb knowledge at them), speaking crudely, midwits with no specialized knowledge currently leave state-of-the-art models weeping in the dust.

Solving Starburst an era earlier than o3-mini can, while presumably only a modest jump in difficulty for humans, would require genuine geometric reasoning. I've thus far seen no signs that LLMs have this ability. I therefore suspect, with low confidence, that o3 and its contemporaries will fail to break this barrier. If a model passes this barrier, that would be a massive cause for concern.

Fig. 1: The earliest era where various models are able to solve Starburst. Era 15 here is past (worse than) where I expect the average person to solve Starburst. Solving it in Era 15 would be a very significant threshold for an LLM to cross.

Read the original on pennheretic.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.