Gemini 3.1 Pro scored comparably to Gemini 3 Pro on FrontierMath.
It also solved a Tier 4 problem that no model has solved before, though not how a human would.
We accidentally ran Gemini 3.1 Pro on Tier 4 a second time. The score above reflects the first, official run. But we noticed in the second run that it had solved a problem no model had solved before. The newly-solved problem is by Emmanuel Breuillard. See his commentary below.
We are waiting for API access to evaluate Gemini 3 Deep Think
Check out our benchmarking hub for more, including Gemini 3.1 Pro scores on other benchmarks!


