In my last post I wrote OpenAI was quiet and might drop something soon against Anthropic. Well, OpenAI didn't just drop something. Everyone dropped.
In literally 48 hours: Gemini 3.8 Flash dropped, Meta dropped Muse Spark 1.3, and now OpenAI dropped GPT-6 Astra. All frontier models. All at once, lol.
Anthropic
- Claude Fable 5.1 and Mythos 5.1 (Sept 1)
- Gemini 3.8 Flash (Sept 2)
Meta
- Muse Spark 1.3 (Sept 2)
OpenAI
- GPT-6 Astra (Sept 3)
OpenAI and others were facing uptime issues just a few hours earlier, and now OpenAI's Astra blog literally broke on launch.
OpenAI made a new graph in 32 hours to include Fable lol:
The Numbers (feels gamed / saturated)
Novel puzzles (ARC-AGI-3) 99.9%. 12.8x Sol (7.8%). This benchmark is basically dead now.
Math (FrontierMath Tier 4 v2) 97.6%. +20% over Sol (83.0%) and +9.8 over Fable 5.1.
Agentic coding (Terminal-Bench 4.0) 57.9%. Fable 5.1 had 55.8% for two days.
Science workflows (Terminal-Bench Science 0.1) 64.6%. 2.9x Sol (22.4%) and +23% over Fable 5.1 (52.6%).
Business workflows (AutomationBench) 41.4%. 2.3x Sol (18.1%) and +32% over Fable 5.1 (31.4%).
3D modeling (BenchCAD) 95.9% vs 84.3% on Fable 5.1.
Clinical (HealthBench Pro) 63.4% vs 58.1% on Fable 5.1.
And the Others
Meta's Muse Spark 1.3 is good too. 75.4% on DeepSWE v1.1 and 88.8% on Terminal-Bench 2.1, tying Sol at 42% lower cost ($0.55/task vs $0.95).
And Google rushed Gemini 3.8 Flash right after 3.7. 73.7% on DeepSWE v1.1.
Thoughts
I haven't even finished testing Fable 5.1 and now I have Astra, Gemini 3.8 Flash, and Muse Spark.
Model shelf life is down to 2 days, lol.
Competition is great I guess.