---
title: "Benchmarking Sonnet 5.5: less thinking for the same answers"
description: "We ran Claude Sonnet 5.5 and Sonnet 5 through the same 48 tasks at every effort"
canonical: https://caveman.so/news/benchmarking-sonnet-5-5
last-updated: 2026-09-28
---

# Benchmarking Sonnet 5.5: less thinking for the same answers

We ran Claude Sonnet 5.5 and Sonnet 5 through the same 48 tasks at every effort
level, with Opus 5.5, Sonnet 4.6 and Haiku 4.5 as reference points. Code graded
every task. That came to 3,268 runs and 5,906 API requests, for $198.35 at list
prices.

At the same effort level, Sonnet 5.5 solved at least as many tasks and usually
spent less. How much less depends on the task. On long reasoning problems, where
Sonnet 5 thinks for a long time, Sonnet 5.5 used a fraction of the tokens. On a
typical task at `high` it used about half the output tokens. On chat, extraction
and long-document work the two cost about the same.

![Headline numbers, Sonnet 5.5 against Sonnet 5 with both at high effort on the same tasks. Output tokens across the suite: 0.25x (95% CI 0.18 to 0.39x, median task 0.49x). Cost per task across the suite: 0.33x ($0.019 against $0.060, median task 0.57x). Tasks solved: +3.8 points (99.6% against 95.8%, CI +0.4 to +8.3). Wall-clock time per task: 0.22x (10.0 s against 45.0 s).](/news/benchmarking-sonnet-5-5/ledger.png)

## What we ran

48 tasks in eight kinds: short chat questions (3), JSON extraction (3), an
executive summary with a word limit (1), exact-answer reasoning problems whose
answers we computed by brute force (21), code generation checked by hidden tests
(10), a question about a 70k-token production log (1), agents that fix small
Python repositories and are graded by hidden tests after the loop ends (6), and
agents that answer questions from an internal wiki (3). No model graded another
model. We checked every grader against a reference solution, which must pass,
and against the untouched starting state, which must fail.

The suite leans on reasoning: 21 of 48 tasks, and 66% of Sonnet 5's cost at
`high`. Keep that in mind when you read the suite-wide numbers below. We also
give the median task, which is less sensitive to the mix.

Sonnet 5.5 and Sonnet 5 ran at off, low, medium, high, xhigh and max: five reps
per task, except three at max on the 11 hardest tasks. "Off" means
`thinking: between_tools` on Sonnet 5.5 and `disabled` on Sonnet 5. The
reference models ran three reps.

## Same effort, less output

With both models at `high`, the default effort on both:

| | Sonnet 5.5 | Sonnet 5 | Suite ratio | Median task |
|---|---|---|---|---|
| Output tokens per task | 1,317 | 5,237 | 0.25× (95% CI 0.18–0.39) | 0.49× |
| Cost per task | $0.0195 | $0.0597 | 0.33× | 0.57× |
| Tasks solved | 99.6% | 95.8% | +3.8 pts (CI +0.4 to +8.3) | |
| Wall time per task | 10.0 s | 45.0 s | 0.22× | |

37 of the 48 tasks were cheaper on Sonnet 5.5. The suite ratio comes out lower
than the median because five long reasoning problems account for most of
Sonnet 5's output. Leaving out the reasoning tasks altogether, Sonnet 5.5
produced 0.56× the output tokens.

The gap grows with effort. At `low` there's no saving on a typical task (median
1.19× the output tokens, 0.99× the cost), and only the hard reasoning problems
come out cheaper. From `medium` up, most tasks get cheaper.

Laid end to end, taking one median-cost run per task, the whole suite at `high`
took Sonnet 5.5 490 seconds and 63k output tokens. Sonnet 5 needed 2,034 seconds
and 232k. In the runs shown, both solved all 48.

## The effort dial moved

Sonnet 5.5 at `high` produced fewer output tokens than Sonnet 5 at `low` (1,317
against 2,724 per task). A setting copied over from Sonnet 5 buys something
different now.

The pairing we would start from: **Sonnet 5 at high → Sonnet 5.5 at medium.**
On this suite, tasks solved went from 95.8% to 99.2%. Cost fell to 0.41× on the
median task and 0.27× across the suite, and 41 of 48 tasks came out cheaper.

![The effort dial, mean per task across 48 tasks. Left: output tokens per task on a log scale at off, low, medium, high, xhigh and max; the Sonnet 5.5 line sits below the Sonnet 5 line at every level, closest at max. Right: tasks solved with 95% confidence bars; Sonnet 5.5 at medium solved 99% of tasks, Sonnet 5 at high solved 96% at 3.69x the cost, and both drop at max.](/news/benchmarking-sonnet-5-5/effort.png)

## On single requests, the difference is thinking

On single-request tasks, the gap is hidden thinking. At `medium`, 83% of
Sonnet 5's output tokens were thinking, against 37% on Sonnet 5.5. At `high` it
was 85% against 47%. We estimate this by subtracting a `count_tokens` count of
the visible text and tool calls from the billed output tokens. The same probe
gives 1% for Sonnet 5 with thinking disabled, as it should.

Agents are different. There, output per task was about the same (0.98× on coding
agents, 1.15× on research agents at `medium`), and the saving came from fewer
turns ([Agents take fewer turns](#agents-take-fewer-turns)).

## Where the bill moves, and where it doesn't

Sonnet 5.5 output tokens as a multiple of Sonnet 5 at `medium`:

- Exact-answer reasoning (21 tasks): **0.19×**
- Code generation (10 tasks): **0.45×**
- Chat, extraction, summary (7 tasks): 1.01–1.14×
- The 70k-token log question (1 task): output 0.13×, but **cost 0.99×**, because
  input dominates and input tokens cost the same on both models

![Sonnet 5.5 output tokens as a multiple of Sonnet 5 by workload, both at medium effort. Chat 1.01x, extraction (JSON) 1.14x, summarization 1.10x, exact-answer reasoning 0.19x, code generation 0.45x, long context of about 70k input tokens 0.13x, wiki research agent 1.15x, coding agent with tests 0.98x. Change in tasks solved per workload ranges from +0 to +7 points.](/news/benchmarking-sonnet-5-5/work.png)

So the saving depends on the mix. Taking this suite's per-category results for
the switch from Sonnet 5 at high to Sonnet 5.5 at medium:

| Mix | Cost vs Sonnet 5 |
|---|---|
| Support / chat app | 98% |
| Document pipeline (long inputs) | 98% |
| Coding agent | 46% |
| Analysis / reasoning | 16% |
| All 48 tasks | 27% |

These mixes are weighted averages of our tasks, not a forecast, and some
categories rest on one to three tasks. Caveman shows the real mix of your own
traffic by workflow, and that's the number to use.

![What the switch from Sonnet 5 at high to Sonnet 5.5 at medium does to a bill, for task mixes built from this suite, as a share of the Sonnet 5 bill. Support or chat app 98%, document pipeline with long inputs 98%, coding agent 46%, analysis or reasoning 16%, everything in this suite 27%. Tasks solved stay at 100% for chat and documents and rise for the other mixes.](/news/benchmarking-sonnet-5-5/bill.png)

## Agents take fewer turns

On the nine tool-using tasks, Sonnet 5.5 at medium finished in 3.9 API requests
on average, against 5.2 for Sonnet 5. Agent cost per task was 0.77×. Every agent
turn re-sends the conversation, so fewer turns cut input spend even when output
stays flat.

Small prompts also start caching sooner. Sonnet 5.5's minimum cacheable prompt
is 512 tokens; Sonnet 5's is 1,024. Our pricing-bug agent starts with a system
prompt and tool list of about 760 tokens. On Sonnet 5 the first two turns sat
below the minimum and billed 1,708 uncached input tokens before caching took
over. On Sonnet 5.5, caching applied from the first turn, and each run billed 12
or fewer uncached input tokens.

![Agent loops on the nine tool-using tasks, at each effort level from off to max. Left: API requests per agent task; Sonnet 5.5 needs fewer requests than Sonnet 5 at every level, 3.9 against 5.2 at medium. Right: share of input tokens served from cache, between about 60% and 78% for both models.](/news/benchmarking-sonnet-5-5/agent.png)

## Waiting

Median time to the first visible token across single-request runs was 0.85 s for
Sonnet 5.5 at medium and 2.9 s for Sonnet 5 at medium. At `high` it was 2.4 s
against 6.3 s. Wall time and time to first token were measured on the client
with 8 to 24 requests in flight, so both include queueing.

## Leave max off

`max` cost 6× as much as `xhigh` on Sonnet 5.5 and solved fewer tasks (89%
against 99.6%). Every max failure was a long answer hitting our 64k output cap,
on both models. On this suite, xhigh was the top of the useful range, and medium
or high covered almost everything.

## The reference models

- **Opus 5.5 at medium** solved 98.6% at $0.042 per task, 2.1× the cost of
  Sonnet 5.5 at high. Our tasks don't include the long, hard work where Opus
  should earn its price, so read this as "not needed here", not as a verdict on
  Opus.
- **Haiku 4.5** is the cheapest per task ($0.011) but solved 60%, which works out
  to $0.0176 per *solved* task. Sonnet 5.5 at low came to $0.0164 per solved task
  at 96%.
- **Sonnet 4.6 at high** solved 91.7% at $0.199 per task, 12× Sonnet 5.5 at
  medium.
- **Thinking off** is not quite like for like. Sonnet 5.5 with `between_tools`
  solved 96.2% and Sonnet 5 with thinking disabled solved 87.9%, but
  `between_tools` still returns short progress notes as thinking blocks (about a
  quarter of its output), where Sonnet 5 returns none.

One trap when comparing across generations: Sonnet 5, Sonnet 5.5 and Opus 5.5
share a tokenizer that counts 1.40× more tokens than Sonnet 4.6 on the same
English text (1.15× on code, 1.23× on logs). Per-million-token prices don't
compare directly.

## What to do

1. **Move Sonnet 5 at high to Sonnet 5.5 at medium**, and compare the two on your
   own traffic before you change the default.
2. **Set effort explicitly.** The levels were recalibrated, so a Sonnet 5 setting
   carried over lands somewhere else. At `low`, don't expect a saving on typical
   tasks.
3. **Don't expect a change on long-input work.** Caching and trimming inputs
   matter more there than the model switch.
4. **Leave max off.** If you use xhigh, give it a large `max_tokens`.
5. **Keep agent histories append-only.** Sonnet 5.5 binds thinking blocks to the
   conversation. Editing earlier turns throws that reasoning away, and on
   accounts created since 31 August 2026 it returns an error.

## Limits

These are our 48 tasks, not your traffic, and they lean on reasoning. The coding
and agent tasks were small enough that almost every configuration solved them,
so they show cost differences better than quality differences. The gains in
tasks solved at medium, high and xhigh are small; their confidence intervals
exclude zero, while the one at low does not. `max` was capped at 64k output
tokens. No refusals occurred. Eleven runs failed with API overload errors; we
retried them, and they don't count against either model. These are measured
benchmark numbers, and they only become savings on your account when your own
traffic shows them.
