Model Releases Hacker News (GPT)

GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index

GPT-6 AstraArtificial Analysisbenchmarktoken efficiency

Artificial Analysis benchmarked GPT-6 Astra across its two flagship indices, with different stories in each. The model's pricing is 2.5x GPT-5.6 Sol's across the board, rising from $4/$20 to $10/$50 per million input/output tokens, retaining a 90% cache-read discount and a 25% cache-write premium. In the Coding Agent Index, the gains look strong; in the Intelligence Index, the price increase weakens the value proposition.

On the Artificial Analysis Coding Agent Index, GPT-6 Astra scores 67 in the Codex harness, roughly matching Claude Opus 5 and Fable 5 in Claude Code and Muse Spark 1.3 in Muse Code; Fable 5.1 in Claude Code still leads with 70. The model is about 70% more token-efficient than GPT-5.6 Sol, using one third the tokens in the Codex harness and one fifth the tokens of Claude Opus 5 (xhigh), with various effort levels sitting on the Pareto frontier. At max effort, GPT-6 Astra costs about the same as GPT-5.6 Sol while scoring 2 points higher, and per task it costs less than half as much as Claude Fable 5 for the same score.

On the Artificial Analysis Intelligence Index, GPT-6 Astra scores 61, equal to GPT-5.6 Sol but 5 points below Claude Fable 5.1 (max with fallback) and also trailing Muse Spark 1.3 (max). It uses ~10% fewer output tokens per task at max effort than GPT-5.6 Sol and defines a new Pareto frontier, yet the 2.5x price increase makes it 75% more expensive per task than its predecessor at max effort. In AA-Omniscience, GPT-6 Astra cuts hallucinations from 92% to 51% at max effort while simultaneously raising accuracy by 4 points. It also improves about 80 Elo points on AA-Briefcase, the long-horizon knowledge-work evaluation involving multi-week projects, many linked tasks, and thousands of source files, with gains in both rubric scores and Analytical Quality.

Read original →

← Back to home