Claude Sonnet 5.5 is here. At the same price as Sonnet 5, it is over 30% faster and rivals Opus 5.5.

2026-09-29
58min read
Updated: 2026-09-30
claude-sonnet-55.webp

Table of Contents

Hello. On September 28, 2026 (US time), Anthropic announced "Claude Sonnet 5.5." Following last week's Opus 5.5, it is the second model in the Claude 5.5 family.

The official announcement opens by introducing it as a clear upgrade from Sonnet 5, running more than 30% faster and up to 30% cheaper for many workloads. The pricing table itself remains unchanged from Sonnet 5. The explanation is that by completing the same tasks using fewer tokens at the same unit price, it ends up being cheaper as a result.

Its positioning is also clear-cut. Complex work requiring careful judgment goes to Opus 5.5, while well-scoped daily tasks, bug fixes, and creating documents, slides, and spreadsheets go to Sonnet 5.5. Officially, Sonnet 5.5 is positioned as a "faster, cheaper" model that complements Opus 5.5.

Based on the official announcement and posts on X from the official Claude account, this article breaks down what has changed with Sonnet 5.5 and how best to use it alongside Opus 5.5.

Key Points of This Announcement

The official announcement highlights improvements over Sonnet 5 across five main areas:

  1. Performance - Terminal-Bench 4.0 jumped from 10.3% on Sonnet 5 to 70.6%. On GDPval-AA, it stands virtually neck-and-neck with Opus 5.5. It is also the first Sonnet model to beat Pokémon Red relying solely on screenshots.
  2. Collaboration - Like Opus 5.5, its writing is clearer and more straightforward than the previous generation. Taking advantage of its speed, it is well-suited for rapidly iterating on moderately complex tasks.
  3. Cost - Pricing remains identical to Sonnet 5. The tokens required for the same tasks have decreased significantly, making it up to 30% cheaper per task.
  4. Speed - Output generation is over 30% faster than Sonnet 5, making it the fastest Sonnet to date.
  5. Alignment & Safety - In automated behavioral audits, almost all metrics are equal to or better than Sonnet 5. It is the first Sonnet model released with cybersecurity safeguards comparable to higher-tier models.

Additionally, the final member of the family—Claude Haiku 5.5, designed for high-volume processing and cost-sensitive use cases—is slated to arrive in the coming weeks.

Benchmarks Show a Massive Leap from Sonnet 5, Nearing Opus 5.5

Here is the comparison table published in the official announcement:

Benchmark comparison table for Sonnet 5.5, Sonnet 5, Opus 5.5, and GPT-6 Sol. Terminal-Bench 4.0: Sonnet 5.5 at 70.6%, GDPval-AA v2.1 at 1844, OSWorld 2.1 at 80.1% (Source: Introducing Claude Sonnet 5.5 \ Anthropic)

Extracting the key metrics yields the following:

BenchmarkSonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Terminal-Bench 4.0 (Terminal agentic coding)70.6%10.3%66.4%Not listed
FrontierCode 1.1 Main (Agentic coding)46.2% (Max) / 52.1% (Xhigh)42.4%54.4%49.3%
CursorBench 4.0 (Agentic coding)55.5%34.1%57.8%Not listed
GDPval-AA v2.1 (Knowledge work, Elo)1844144918461487
AA-Briefcase v1.1 (Long-horizon knowledge work, Elo)1811135918221483
Humanity's Last Exam (Cross-domain reasoning, with tools)64.5%54.9%67.7%Not listed
OSWorld 2.1 (Computer use, partial)80.1%57.0%81.8%Not listed
Chartography (Chart interpretation, without tools)61.6%15.6%64.4%53.6%

While every category shows a substantial leap from Sonnet 5, Terminal-Bench 4.0 stands out the most. Jumping by an order of magnitude from Sonnet 5’s 10.3% to 70.6%, it actually surpasses Opus 5.5’s 66.4% on paper. Chartography, which measures chart reading capabilities, also rose from 15.6% to 61.6%, showing a massive improvement in image understanding.

On GDPval-AA v2.1 for knowledge work, it scored 1844—a mere 2-point difference from Opus 5.5's 1846. GDPval-AA evaluates practical tasks across 44 occupations and 9 major industries; Sonnet 5.5 gained roughly 400 points over Sonnet 5.

On the other hand, Opus 5.5 still leads in every category other than Terminal-Bench 4.0. The official announcement explicitly notes that benchmarks capture only one aspect of a model's capabilities, and that Opus 5.5 remains distinctly stronger for complex, open-ended tasks requiring sustained judgment. It is also worth keeping in mind that these figures were published by Anthropic and have not been independently verified by a third party.

Points to Note When Reading the Numbers

The footnotes contain several important caveats regarding the figures:

  • The Opus 5.5 score for Terminal-Bench 4.0 reflects the Xhigh effort setting, which yielded its highest score.
  • FrontierCode evaluates whether changes can be merged cleanly without human intervention, penalizing out-of-scope edits even if they are high quality. Sonnet 5.5 scored lower on Max effort than on Xhigh; this is explained by Max more frequently triggering Claude Code’s code-review skill (splitting reviews across numerous sub-agents), leading to timeouts or extra out-of-scope modifications.
  • GDPval-AA and AA-Briefcase were evaluated by Artificial Analysis using a pre-release version of Sonnet 5.5 that had a bug degrading responses for requests using structured outputs (since fixed). Any impact is considered minor and would have skewed scores slightly downward.

The FrontierCode finding serves as a useful reference when dialing up effort in Claude Code. Increasing effort causes the model to work more thoroughly, but it may also broaden its focus beyond the requested scope. As noted in the Opus 5.5 article—where Terminal-Bench 4.0 peaked at xhigh and dipped slightly at max—these footnotes reaffirm that "higher effort isn't always strictly better."

Outperforming Sonnet 5’s Peak Scores Even at Lower Effort

This is perhaps the most practically impactful takeaway of the announcement. According to Anthropic, Sonnet 5.5 at Low or Medium effort outperforms Sonnet 5’s highest scores across several benchmarks at roughly one-tenth the cost per task.

Effort is a setting that determines how deeply the model deliberates before responding, spanning five levels: Low, Medium, High, Xhigh, and Max. Higher settings allow for deeper thinking at the expense of tokens and time, while lower settings are faster and cheaper. In the official graphs, the horizontal axis shows cost per task (logarithmic scale) and the vertical axis shows score, with points connected across effort levels. Points toward the top-left represent "cheaper and higher-scoring."

Terminal-Bench 4.0 scores and costs by effort level. Sonnet 5.5 ranges from ~20% at Low to 70.6% at Max, while the Sonnet 5 line remains around 10% (Source: Introducing Claude Sonnet 5.5 \ Anthropic)

Looking at the Terminal-Bench 4.0 chart, Sonnet 5 stays around 10% (roughly $12 per run) even at its highest effort. In contrast, Sonnet 5.5 reaches roughly 29% (~$0.8) at Medium, which is the default in Claude apps. True to the official claim, it significantly beats Sonnet 5's top score at less than one-tenth the cost.

However, its relationship with Opus 5.5 varies by effort level. Up through roughly High effort, Opus 5.5 maintains a higher curve; Sonnet 5.5 achieved its 70.6% score at Max ($13). Opus 5.5 achieved 66.4% at Xhigh ($7.5), meaning that while Sonnet 5.5 did edge out Opus 5.5 on Terminal-Bench 4.0, it incurred a proportional cost to do so.

AA-Briefcase v1.1 Elo and costs by effort level. Sonnet 5.5 at Medium beats Sonnet 5's top score, tracking nearly identical to Opus 5.5 at High and above (Source: Introducing Claude Sonnet 5.5 \ Anthropic)

AA-Briefcase v1.1 is a new benchmark introduced by Artificial Analysis to measure sustained, long-horizon knowledge work. Sonnet 5.5 at Medium scored roughly 1460 (~$1.6), beating Sonnet 5’s peak score (1359 at ~$14) at about one-ninth the cost. From High effort onward, the curves for Sonnet 5.5 and Opus 5.5 overlap almost completely, which is quite striking.

The shape of this chart provides a clear rule of thumb for workflow division. Anthropic explains that Sonnet 5.5 best complements Opus 5.5 when run at lower effort levels, though at higher effort levels it can produce comparable results at similar costs. In short, a clean separation of roles emerges: use Sonnet 5.5 at Low or Medium for fast, cost-effective iteration, and save Opus 5.5 for tasks requiring deep, deliberate contemplation.

Note: The chart figures are approximations read from the axes; the measurements themselves were conducted by Anthropic. Additionally, the announcement notes that on FrontierCode, Sonnet 5.5 at High effort (Claude Platform's default) matches GPT-6 Sol’s peak score at approximately one-fifth the cost.

Default Effort Is Medium for Claude Code and Apps

According to Anthropic, the default effort is Medium for Claude Code and the Claude app, and High for the Claude Platform (API). Lower settings are faster and use fewer tokens, making them ideal for routine work, while higher settings deliberate longer and check their work more thoroughly.

Fast Codebase Comprehension in Coding Tasks

Anthropic highlights coding as an area with particularly noticeable gains. On FrontierCode, comparing both models at High effort, Sonnet 5.5 scored 10 points higher than Sonnet 5 at roughly one-fifteenth the cost per task. On CursorBench, which evaluates tasks drawn from real-world Cursor sessions, its peak score sits within roughly 2 points of Opus 5.5.

A common thread across early tester feedback is rapid codebase understanding and high operational efficiency. Direct comparisons showed it tends to batch tool calls together more effectively than Sonnet 5, reducing total steps and cost. Here is a summary of notable tester remarks:

  • Epic Games noted that it met the quality standards expected of top-tier models, successfully handling gameplay systems engineering tasks involving tens of thousands of lines of code spanning multiple hours.
  • CodeRabbit stated that the tendencies to over-rely on web search and consume excessive tokens—both pain points with Sonnet 5—were resolved. They plan to transition simple to medium-complexity reviews to Sonnet 5.5 first.
  • Base44 reported building applications with quality comparable to Opus 5 across 118 real-world application builds. It required an average of 3.6 iterations per build, less than half of Opus 5’s 7.7 iterations.
  • Lovable observed that in internal coding evaluations, tool calls required to complete tasks decreased by one-third, and shell execution counts were roughly halved.

Another quote from game creator Kevin Ngo neatly illustrates how to split tasks with Opus 5.5:

When Claude Opus 5.5 sets the architecture and general framework for a game, I would feel confident in letting Sonnet 5.5 implement it.

The takeaway is that once Opus 5.5 handles the architecture and overarching framework, implementation can be safely delegated to Sonnet 5.5. In Claude Code, this points toward a clear pattern: using Opus 5.5 for planning, while dispatching implementation and minor tweaks to Sonnet 5.5 sub-agents.

Praised for Design Sensibility in Knowledge Work

In knowledge work, Sonnet 5.5 closely approaches Opus 5.5 in computer use and chart interpretation, while clearly outperforming Sonnet 5 and GPT-6 Sol in long-horizon knowledge tasks.

As an improvement harder to capture in raw numbers, Anthropic points to its design sensibility. It refines user interfaces and builds presentations aligned to slide templates that require almost no manual revision. In an internal test, it was provided with a public company's quarterly earnings materials, presentation transcripts, and a slide template to produce a 10-slide performance review. Two domain experts deemed the first draft ready to send as-is.

On X, a side-by-side video demonstrated Sonnet 5 and Sonnet 5.5 generating code for the prompt: "simulate wind shaping sand dunes in a single HTML file." The official page also features comparisons generating a flock of 400 starlings and a clock composed of 24 smaller clocks.

A few reports from early testers:

  • Slack reported that without changing prompts, Sonnet 5.5 outperformed Sonnet 5 across almost all offline Slackbot evaluations, requiring fewer steps and roughly 14% fewer output tokens.
  • Box noted that it cross-referenced source documents to verify data, catching errors that Sonnet 5 missed. It was more accurate than its predecessor, 2.4x faster, and consumed 12% fewer total tokens.
  • Investment firm Balyasny Asset Management found it outperformed Sonnet 5 across 2,441 financial tasks, using roughly 121k tokens per response compared to Sonnet 5's ~497k tokens.

The reduction in token volume stands out across these partner reports. Even with identical rate cards, actual billing may end up substantially lower.

Identical Pricing to Sonnet 5, Half the Cost of Opus 5.5 for I/O

Pricing is as follows:

Metric (per 1M tokens)Sonnet 5.5Opus 5.5
Input$2$4
Output$10$20
Cache Read$0.20$0.20
Cache Write$2.50$5

Input and output are exactly half the price of Opus 5.5. Cache reads (reading previously processed input) match at $0.20 following Opus 5.5's price reduction.

Because it charges the same unit price as Sonnet 5 while consuming fewer tokens per task, official testing indicates it is up to 30% cheaper per task. Output generation is also over 30% faster, a difference in efficiency that is immediately apparent in the official comparison videos.

Safety and Alignment

The safety section opens with the caveat that "Sonnet 5.5 does not advance the state of the art in frontier capabilities." Consequently, alignment evaluations focused on risks that apply regardless of capability tier: acting against user interests, misleading users, and assisting in severe misuse.

Key evaluation findings include:

  • Automated Behavioral Audits - In audits evaluating behavior across approximately 1,850 scenarios, it matched or exceeded Sonnet 5 across nearly all metrics for alignment, misuse resistance, and honesty.
  • Containment Evaluations - Its low rate of sandbox escape attempts approached that of Opus 5.5 (the top performer tested), and its tendency to probe container boundaries was the lowest among all Anthropic models.
  • Overall - While Opus 5.5 performed slightly better across the broader audit, no evidence was found of Sonnet 5.5 pursuing goals in conflict with user intent.

At the same time, Anthropic notes that no evaluation suite catches every failure mode, and Sonnet 5.5 may harbor yet-undiscovered behaviors. For this reason, it is paired with the following safeguards.

Cybersecurity Safeguards Elevated to Opus 5.5 Standards

Sonnet 5.5 incorporates safeguards across three main areas. To clarify, Sonnet previously had safeguards in place: in the June Sonnet 5 announcement, Sonnet 5 launched with the same cyber safeguards as Opus 4.7 and 4.8 enabled by default (a looser configuration than Fable 5, as cyber risks were assessed to be low). With Sonnet 5.5, these safeguards have been elevated to match the level of Opus 5.5.

  • Cybersecurity - Because its cyber capabilities have improved substantially over Sonnet 5 to rival Opus 5, it is released with safeguards equivalent to Opus 5.5. While it can identify and fix code bugs during standard development, high-risk cybersecurity tasks will transparently fall back to Sonnet 5. Defensive practitioners will soon be able to apply for an expanded Cyber Verification Program, granting tiered access to advanced capabilities across Sonnet 5.5, Opus 5.5, and Mythos models.
  • Biology - Biological safeguards remain unchanged from Sonnet 5. Designed to target harmful requests, they do not interfere with most research, educational, or clinical work, though some requests in microbiology and virology may trigger false positives. Organizations requiring broad biological workflows can apply for the Life Sciences Verification Program.
  • Distillation Resistance - To defend against distillation attacks using large numbers of synthetic accounts to siphon model capabilities, it is the first Sonnet model released with safety classifiers designed to prevent reasoning extraction. Additionally, the scope of preserved thinking has expanded, permanently binding Claude’s thinking tokens to the originating account.

Distillation defenses are expected to be transparent to most developers. However, workflows that transfer conversations across accounts—such as switching accounts mid-session in Claude Code—will be affected. If this applies to your workflow, reviewing the documentation linked on the official page is recommended.

For Claude Code users, the primary point of note will be the fallback to Sonnet 5 for certain security-related operations. While standard bug fixing should be unaffected, tasks like vulnerability assessments deemed high-risk may behave differently than before.

API Migration: Watch the Thinking Configuration

Like previous Sonnet releases, Sonnet 5.5 supports zero data retention. Availability began on launch day across all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. The API model ID is claude-sonnet-5-5.

When migrating from Sonnet 5 on the API, special attention is required if your implementation disables thinking (reasoning before answering). According to the migration guide, using thinking: {"type": "disabled"}—the syntax used to disable thinking in Sonnet 5—will return a 400 error on Sonnet 5.5. Instead, you must specify the newly added between_tools.

{
  "model": "claude-sonnet-5-5",
  "max_tokens": 16000,
  "thinking": {"type": "between_tools"},
  "output_config": {"effort": "high"},
  "messages": [{"role": "user", "content": "..."}]
}

between_tools is the lowest thinking tier in Sonnet 5.5; it bypasses deliberation prior to the initial response, but thinking blocks will still be returned between tool calls. The migration guide notes several additional constraints:

  • Supported effort levels are limited to low, medium, and high; pairing this mode with xhigh or max results in a 400 error.
  • Effort levels cannot be modified mid-conversation. To vary effort turn-by-turn, use adaptive thinking.
  • Older SDKs that lack the between_tools definition may fail type checking in Python and TypeScript. Update the SDK or pass the payload as raw JSON.

Furthermore, Sonnet 5.5 enables thinking by default if thinking is omitted. If your existing code ran without thinking, it is wise to review the following:

  • Responses may begin with a thinking block; inspect the block's type rather than assuming content[0].text.
  • In tool call loops, pass thinking blocks back unmodified, including empty ones.
  • max_tokens includes thinking tokens, and thinking tokens are billed as output tokens, so upper limits may need adjustment.

The migration guide also details using Claude Code’s /claude-api migrate command to automate migration via its built-in Claude API skill. It handles model ID replacement, fixes invalid parameters, adjusts effort levels, and outputs a checklist of items requiring manual verification.

Trying It Out in Claude Code

Checking my local Claude Code environment (v2.1.284), the model alias sonnet already pointed to claude-sonnet-5-5. You can launch it using either of the following commands:

claude --model sonnet
claude --model claude-sonnet-5-5

As illustrated in the earlier graphs, Sonnet 5.5 shines in its cost-efficiency at lower effort levels. A great starting point is to leave it at the default Medium setting and assign it daily bug fixes or minor feature implementations.

If your environment does not default to Sonnet 5.5, check your version with claude --version, update if necessary, and restart:

claude update

Model availability and default behaviors vary by plan, provider, and admin configuration. If you are still running Claude Code via Homebrew or npm, transitioning to a native installation will enable auto-updates and ensure smoother adoption of new models. Detailed steps are covered in How Switching Claude Code from Homebrew to Native Install Improved My Workflow.

Comparing Sonnet 5 and Sonnet 5.5

A comparison of key differences:

ItemSonnet 5Sonnet 5.5
Pricing (Input / Output per 1M tokens)$2 / $10$2 / $10 (Unchanged)
Cost per TaskBaselineUp to 30% reduction
Output Generation SpeedBaseline30%+ faster
Terminal-Bench 4.010.3%70.6%
CursorBench 4.034.1%55.5%
GDPval-AA v2.114491844
Cybersecurity SafeguardsIdentical to Opus 4.7/4.8 (Looser than Fable 5)Equivalent to Opus 5.5 (Interventions fall back to Sonnet 5)
Disable Thinking Syntaxdisabledbetween_tools
API Model IDclaude-sonnet-5claude-sonnet-5-5

Benchmark figures are based on the official announcement comparison table.

What This Means for Users

Recapping what these changes mean in practical day-to-day work:

1. Significantly Higher Performance at the Same Price

Terminal-Bench 4.0 jumped from 10.3% to 70.6%, and GDPval-AA is virtually level with Opus 5.5. The rate card remains identical to Sonnet 5, meaning users on default configurations can upgrade simply by updating the model ID. API implementations disabling thinking will require the syntax update mentioned earlier.

2. Highly Capable Even at Low Effort

Sonnet 5.5 at Low or Medium outperforms Sonnet 5’s peak scores at roughly one-tenth the cost. This makes a noticeable difference for repetitive workloads, such as sub-agents and high-volume batch processing.

3. Clearer Division of Labor with Opus 5.5

Because high-effort Sonnet 5.5 approaches Opus 5.5 performance at roughly comparable costs on some benchmarks, roles are easier to delineate: use Opus 5.5 when you need deep reasoning, and use Sonnet 5.5 at lower effort when you need speed and cost efficiency.

4. Cyber Workflows and API Thinking Settings Require Review

Points to keep in mind: high-risk security tasks will fall back to Sonnet 5, and API code using thinking: {"type": "disabled"} must be updated to between_tools.

Summary

Key takeaways for Claude Sonnet 5.5:

  • Second model in the Claude 5.5 family. Positioned as a faster, cheaper complement to Opus 5.5. Haiku 5.5 is scheduled to follow in the coming weeks.
  • Scores 70.6% on Terminal-Bench 4.0 (vs. Sonnet 5's 10.3%, Opus 5.5's 66.4%), 55.5% on CursorBench 4.0, and 1844 on GDPval-AA v2.1 (vs. Opus 5.5's 1846).
  • Opus 5.5 leads across benchmarks outside Terminal-Bench 4.0, maintaining a clear advantage in complex, open-ended tasks.
  • Outperforms Sonnet 5's peak scores even at Low or Medium effort on several benchmarks at roughly one-tenth the cost.
  • Pricing matches Sonnet 5: $2 input, $10 output, and $0.20 cache read (per 1M tokens). Up to 30% cheaper per task, with 30%+ faster output generation.
  • Default effort is Medium for Claude Code and consumer apps, High for Claude Platform.
  • Improved design sensibility across UI layouts and slide presentation polish.
  • Matches or exceeds Sonnet 5 on nearly all automated behavioral audit metrics.
  • Cyber safeguards raised to Opus 5.5 standards; first Sonnet model equipped with distillation-defense classifiers. High-risk security queries fall back to Sonnet 5.
  • API model ID is claude-sonnet-5-5. To bypass pre-response reasoning, use between_tools instead of disabled.

In my Opus 5.5 post, I noted the trend where capabilities pioneered at the top tier trickle down to accessible price points within weeks. This time, that transition reached Sonnet in about a week. With a model that rivals Opus 5.5 on select knowledge work benchmarks at half the input/output cost, it offers a great opportunity to rethink how tasks are distributed across models.

For now, I plan to keep Opus 5.5 as my primary model in Claude Code while offloading minor fixes and sub-agent tasks to Sonnet 5.5 to get a feel for the workflow. I'll share further observations in a future post.

Reference Links

Share this article

Related Articles