RSS Amplifier

Stephen's Exhaust Pipe · Nov 18, 2025

Last minute Gemini 3 predictions

0
Sign in to vote or save

Stephen Malina · Stephen's Exhaust Pipe

As Google has not so subtly been hinting at, Gemini 3 is likely coming out tomorrow. As with GPT-5, I thought it could be fun to pre-register some predictions about Gemini 3 as a way to ground myself and test my Delphic mettle.

Overall, my mental model here is that Gemini will be a big jump for Google but not a step change relative to the current SoTA models (GPT-5.1 Thinking, Claude 4.5 Sonnet, etc.). My benchmark predictions basically align with that, treating slightly better than whichever of the current SoTA models are on top as reasonably likely, but off trend improvements as much less likely.

Many of the below predictions are on static benchmarks. Per my GPT-5 post, I think the value of predicting performance on this type of benchmark continues to decline over time. But unfortunately, they remain the easiest thing to make predictions for (in large part thanks to this amazing Epoch AI resource), so given my time constraints, I ended up focusing on them here.

Note: If Google releases multiple thinking levels, I’ll default to the highest level benchmarked for a given benchmark.

By default, my source for scores will be the announcements post unless another source is linked for a given benchmark below.

  • METR time horizon (50% success):

    • >2 hrs: 65%

    • >2.5 hrs: 55%

    • >3 hrs: 35%

    • >3.5 hrs: 5%

  • GPQA Diamond score (pass@1):

    • >80%: 95%

    • >85%: 70%

    • >90%: 40%

  • FrontierMath score (pass@1):

    • Overall:

      • >30%: 90%

      • >40%: 60%

      • >50%: 40%

      • >70%: 20%

    • On Tier 4:

      • >10%: 60%

      • >20%: 40%

      • >30%: 30%

  • SWE-Bench Verified score (pass@1):

    • >80%: 40%

    • >85%: 25%

    • >90%: 5%

  • Terminal Bench score (pass@1):

    • >50%: 45%

    • >60%: 35%

    • >80%: 5%

  • OSWorld score:

    • >60%: 60%

    • >70%: 30%

    • >80%: 15%

  • BALROG score:

    • >40%: 95%

    • >50%: 40%

    • >60%: 10%

  • NYT Connections Benchmark score:

    • >60%: 70%

    • >90%: 25%

    • >95%: 5%

  • Gemini 3 unlocks meaningful new agentic use-case that has been blocked until now (e.g. widespread flight booking): 40%

  • Gemini 3 includes video out as part of the main model: 30%

  • Gemini 3 includes agentic tool usage as part of its thinking: 80%

  • Gemini 3 wows me with its intelligence relative to current SoTA models after 3 hours of cumulative use: 35%

  • I fail to find clear smells or issues in Gemini 3’s code after using it to write 5 nontrivial pull request size changes: 25%

Read the original on an1lam.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.