GPT-6 Astra Benchmark Results
OpenAI claims GPT-6 Astra is the world's most intelligent AI model, but independent testing tells a different story. While Astra leads in specific areas like computer automation and cybersecurity, it falls behind competitors in general intelligence benchmarks.
Key Performance Comparisons
General Intelligence
Artificial Analysis's Intelligence Index scores Astra at 61, the same as its predecessor GPT-5.6 Sol. This puts it behind Claude Fable 5.1 (66) and Meta's Muse Spark 1.3. The model shows no aggregate improvement in broad reasoning capabilities despite its higher cost.
Coding Capabilities
Astra performs well on terminal-based coding tasks, scoring 57.9% on Terminal Bench 4.0 versus 37.3% for GPT-5.6 Sol. However, on general coding benchmarks like Deep SWE, it barely edges out competitors with a 74.1% score compared to 73.8% for Gemini Flash.
Specialized Strengths
Astra shows clear advantages in:
- Computer automation (OS World 2.0 score of 72.6%)
- Cybersecurity (100% on ExploitBench)
- Long-context retrieval (96% accuracy on million-token tests)
Cost and Efficiency
Astra's API pricing is $10/$50 per million tokens versus $4/$20 for GPT-5.6 Sol. While more token-efficient, total task costs run about 75% higher at maximum effort. Some benchmarks revealed regressions in banking tool use and document reasoning.
Benchmark Controversies
The model's 99.9% ARC-AGI-3 score used OpenAI's proprietary testing setup. Under standard conditions, it scored 62.7%. This makes direct comparisons with previous models difficult as they were tested differently.
Practical Implications
Businesses considering Astra should evaluate their specific needs:
- Worth the premium for automation and security tasks
- Less compelling for general AI applications 1.5x higher cost than GPT-5.6 with limited broad improvements
