

Date Published
December 12, 2025
Total Read
3 min
Tags
OpenAI released GPT-5.2 yesterday, marking their 10th anniversary with a model that achieves record-breaking results across numerous benchmarks. But beneath the impressive numbers lies a more nuanced story about how we measure AI progress and where it is taking us.
GPT-5.2 Thinking claims to be the first model performing at or above human expert level on GDPVal, beating top professionals on 71% of comparisons. However, several critical caveats deserve attention:
The benchmark focuses on well-specified digital tasks with full context provided upfront
Real-world tasks often require tacit knowledge and contextual discovery
Catastrophic mistakes (like deleting entire hard drives) are not factored into the scoring
Only 44 occupations were tested, excluding non-digital work
The model excels at tasks like web research and spreadsheet creation. When asked to create a football-themed interaction matrix with match results, GPT-5.2 Pro delivered accurate results with impressive formatting. But the base GPT-5.2 (what most users access) could not complete the same task successfully.
Performance increasingly depends on thinking time and token spend. More tokens generally mean better results, making model comparisons increasingly difficult. On ArcAGI-1, GPT-5.2 Pro Extra High Reasoning effort achieves over 90%, but this requires significant computational resources.
Noam Brown from OpenAI admits that single-number benchmark results are oversimplified. Ideally, evaluations would show performance relative to cost or token usage on the x-axis, revealing the actual trade-offs.
While OpenAI previously showed intellectual honesty by comparing GPT-5.0 High against Claude Opus 4.1 (which performed better), they have not compared GPT-5.2 against Claude Opus 4.5 or Gemini 3 Pro. This has led community members to make their own comparisons, such as Logan Kilpatrick demonstrating Gemini 3 Pro's superior image segmentation on the same motherboard example that OpenAI showcased. (Source)

Different benchmarks testing ostensibly the same capabilities give conflicting results:
MMMU Pro
(table/chart analysis): Gemini 3 Pro leads at 81% vs GPT-5.2's 80.4%
Charkive reasoning
(chart understanding): GPT-5.2 wins at 88.7% vs Gemini's 81%
Both claim to test visual reasoning, yet produce opposite conclusions about which model is superior.
On SimpleBench (a private benchmark testing common sense and spatial-temporal reasoning), GPT-5.2 Pro achieved 57.4% compared to the ~84% human baseline. Gemini 3 Pro performed significantly better at 76.4%. The base GPT-5.2 scored 45.8%, slightly below GPT-5.1.
This suggests potential benchmark maximisation, where performance on publicised coding and maths benchmarks might come at the expense of general reasoning capabilities.
GPT-5.2 does show real improvements:
Near 100% accuracy on four-needle retrieval across long contexts (up to 400,000 tokens)
Competitive pricing via API (cheaper than Opus, input tokens cheaper than Gemini 3 Pro)
Improved mental health evaluation performance
Better long-context recall than previous models
For contexts exceeding 400,000 tokens, Gemini 3 (which handles up to 1 million tokens) remains the better choice.
Sam Altman stated that in 10 years, we are "almost certain to build superintelligence." OpenAI has already moved beyond GPT-5.2 to develop its next model.
But the fundamental question remains: is this incremental approach, ticking off human tasks one by one, breaking benchmarks sequentially, the actual path to AGI?
Think of it like counting sheep across vast fields. Each sheep represents a human task we want to automate. LLMs are methodically counting each sheep, field by field. Many hoped for a flash of inspiration that would lead someone to write an algorithm that would instantly scan all fields and count every sheep simultaneously, a one-shot superintelligence.

That flash may never come from humans. But if we continue this incremental progress, one benchmark broken, one human baseline exceeded after another—eventually, we would count all the sheep.
Whether that constitutes accurate general intelligence or merely comprehensive narrow intelligence across many domains remains the central question of our time.