Anthropic's Claude Sonnet 5.5 Is Out, Beats Opus 5.5 at Coding for Half the Price

Anthropic's mid-tier model tops its own flagship on Terminal-Bench 4.0 and costs half as much per token, but an independent tester found it burns more tokens than any model it has measured.

By Jose Antonio Lanz

3 min read

Anthropic released Claude Sonnet 5.5 on Monday, an upgrade to Sonnet 5 from June. Anthropic says this middle-tier model runs more than 30% faster than its predecessor.

“Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It’s also got a sharp eye for design,” Anthropic wrote.

Myriad: How low will Nvidia go? Click to make your prediction.

The price stays at $2 per million input tokens and $10 per million output tokens. Tokens are the chunks of text an AI reads and writes, a bit shorter than a word, and companies bill by the million. That is half what Opus 5.5 charges. However, it uses nearly a third less tokens per task, which means it ends up being cheaper to run than Sonnet 5.

That said, this model shines in coding. On Terminal-Bench 4.0—a test of whether an AI agent can finish complex professional tasks by typing commands on its own, scored as the share of tasks completed—Sonnet 5.5 hit 70.6%. Opus 5.5 scored 66.4%, and Sonnet 5 managed 10.3%.

In plain terms, the cheaper model finished more jobs. Artificial Analysis, an independent testing firm, ran its own version and agrees: 63.6% for Sonnet 5.5, 59.6% for Opus 5.5, and 59.1% for OpenAI's GPT-6 Astra.

Scores also depend on the effort setting, a dial that makes a model think longer for a better answer and a bigger bill. Anthropic says Sonnet 5.5 at High effort matches GPT-6 Sol on FrontierCode for about a fifth of the cost per task.

On GDPval-AA, which grades real-world professional work across 44 occupations using Elo—the chess-style system that ranks relative skill—Sonnet 5.5 scored 1844 to Opus 5.5's 1846, effectively a tie. GPT-6 Sol scored 1487.

Rivals match the price. OpenAI cut GPT-6 Sol to $2 and $10 last week, and GPT-5.6 Terra, its mid-tier model, lists at $2 and $12. Anthropic published no Terra benchmarks.

The catch

Sonnet 5.5 is a heavy talker. At max effort it wrote about 193,000 tokens per test task, the most Artificial Analysis has measured and roughly 60% more than Opus 5.5. That came to $7.60 per task, about 50% above Sonnet 5, which cuts against Anthropic's claim of up to 30% savings.

Anthropic's savings come from lower settings: at Medium effort, the default in its apps, it says Sonnet 5.5 beats Sonnet 5's best coding score for less than a tenth of the cost. Artificial Analysis says High effort is the best value. For everyday users, that means near-flagship coding at a fraction of the price, as long as the dial stays low.

Anthropic's table is self-reported, and Artificial Analysis tested a pre-release build with a bug that Anthropic expects changed little or slightly understated its scores. Anthropic says Opus 5.5 remains clearly stronger at complex work needing sustained judgment.

Claude Haiku 5.5, built for high-volume, cost-sensitive applications, is due in the coming weeks.

Get crypto news straight to your inbox--

sign up for the Decrypt Daily below. (It’s free).

Recommended News