Model Releases Hacker News (Claude)

Claude Fable 5.1 made me a nice animated pelican

Claude Fable 5.1AnthropicTerminal-Bench-Sciencereasoning

Anthropic released Claude Fable 5.1 (and Mythos 5.1) on September 1st, 2026, claiming it "sets a new standard for coding, knowledge work, and long-running problem-solving tasks." The announcement highlighted scientific research, boasting a 52.6% score on the new Terminal-Bench-Science 0.1 benchmark (first announced August 27th), up from 24.7% for Fable 5, 29.0% for Opus 5, and 22.4% for GPT-5.6 Sol. Other benchmarks showed only slight improvements by comparison.

The author revisits their personal "pelican benchmark" (generating an SVG of a pelican riding a bicycle), noting its correlation with other tasks has weakened but it remains useful for comparing within model families and across reasoning effort levels. Fable 5.1 offers five reasoning levels—low, medium, high, xhigh, max—with no option to disable reasoning entirely. After fixing an issue in llm-anthropic that prevented reasoning traces from being recorded, the author ran the same prompt at all five levels and recorded output tokens, time, cost, and reasoning transcripts.

At low and medium effort, the model produced no reasoning text at all—1,998 output tokens in 23.8 seconds costing 10.017 cents, and 1,977 tokens in 23 seconds costing 9.912 cents, respectively. This was surprising, as the model appeared to skip reasoning entirely for this prompt. At high effort, it produced a short reasoning trace (about 2,612 total output tokens, 29.6 seconds, 13.087 cents) that simply planned the SVG layout.

At xhigh effort, the output exploded to 36,767 output tokens, 7 minutes 51 seconds, and $1.83, with a lengthy reasoning trace including details like keeping the pelican intentionally oversized for comic effect and accepting slight thickness as "charming rather than overengineering it." At max effort, the model used 65,927 output tokens over 13 minutes 54 seconds at a cost of $3.30, yielding the author's best pelican yet, with a detailed background. The author highlights stark differences in how each reasoning level handled the same task.

Read original →

← Back to home