Claude Opus 5.5 is here. Fable 5.1-level performance, 40% cheaper than Opus 5

2026-09-23
62min read
Updated: 2026-09-23
claude-opus-55.webp

Table of Contents

Hello. On September 22, 2026 (US time), Anthropic announced "Claude Opus 5.5." It is the first model in the new Claude 5.5 family, with Sonnet 5.5 and Haiku 5.5 expected to follow in the coming weeks.

July's Opus 5 was positioned as "intelligence approaching Fable 5 at Opus prices." Earlier this month, Fable 5.1 was released, putting the Fable side ahead once again. This new Opus 5.5 is positioned to deliver performance on par with Fable 5.1 across many tasks, at an even lower cost than Opus 5.

The opening of the official announcement is largely summarized in this single sentence:

It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

It is a declaration that it operates at the same level as Fable 5.1 for most tasks while costing 40% less to run than Opus 5. In this article, based on the official announcement, I will break down what has changed with Opus 5.5 and where Claude Code users will feel the biggest impact. In my own Claude Code environment, Opus 5.5 has already become the default model.

Key Points of This Announcement

The official announcement highlights four major areas of improvement for Opus 5.5:

  1. Performance - A major leap forward from Opus 5, reaching parity with Fable 5.1 across many tasks. One early tester reportedly completed a 680,000-line code migration in less than a day.
  2. Safety - Achieved the highest scores to date in the company's automated behavioral audits (alignment evaluations verifying behavior across thousands of simulated environments). Irreversible operations and actions outside the assigned scope have been significantly reduced.
  3. Cost and Speed - Under default settings, it is approximately 40% cheaper than Opus 5 on typical workloads, with output generation over 30% faster.
  4. Communication - Prose is easier to read, front-loading critical information.

Additionally, Opus 5.5 marks the first release since Dario Amodei called for "pacing the frontier." Prior to release, it also underwent testing by external evaluation bodies such as METR and Frontier Design. I will cover the safety aspects in a later section.

Outperforming Fable 5.1 and GPT-6 Astra Across Many Benchmark Categories

Here is the comparison table published in the official announcement:

Benchmark comparison table for Opus 5.5, Fable 5.1, Opus 5, GPT-6 Astra, and GPT-5.6 Sol. Terminal-Bench 4.0 is 66.4% for Opus 5.5, CursorBench 4.0 is 57.8%, GDPval-AA v2.1 is 1846 (Source: Introducing Claude Opus 5.5 \ Anthropic)

Extracting the key metrics yields the following:

BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 Astra
Terminal-Bench 4.0 (Terminal-based agentic coding)66.4%55.8%52.3%57.9%
FrontierCode v1.1 (Agentic coding)54.4%50.3%48.0%53.3%
CursorBench 4.0 (Agentic coding)57.8%51.8%46.6%Not listed
GDPval-AA v2.1 (Knowledge work, Elo)1846173517081542
AutomationBench (Business workflows)40.0%31.4%26.9%41.4%
Terminal-Bench-Science 0.1 (Scientific research as an agent)58.7%52.6%29.0%64.6%
OSWorld 2.0 (Computer use, partial)81.8%80.7%74.0%Not listed

The most eye-catching result is Terminal-Bench 4.0. It jumped 14.1 points from Opus 5's 52.3% to 66.4%, surpassing Fable 5.1's 55.8% by more than 10 points. Since I consider this the benchmark closest to how Claude Code is actually used, this improvement is genuinely exciting.

On the other hand, GPT-6 Astra leads in AutomationBench and Terminal-Bench-Science 0.1, showing that it does not lead in every category. It is also worth keeping in mind that the figures in the comparison table were published by Anthropic itself and have not been independently verified by third parties.

Caveats When Reading the Numbers

The footnotes contain several important caveats:

  • Unless noted otherwise, Opus 5.5 results were evaluated using adaptive thinking set to max effort. For Terminal-Bench 4.0 only, the table reports the highest scores achieved: xhigh for Opus 5.5 and high for GPT-6 Astra.
  • Evaluations were conducted with the same safeguards enabled as in actual production use for Opus 5.5 (similar to Fable 5.1). If safeguards intervene during evaluation, the task is handed off to a fallback model to continue. Cyber-related tasks route to Opus 4.8, while biology and frontier LLM development tasks route to Opus 5. Because earlier models solved portions of these tasks instead of Opus 5.5, the reported scores may skew slightly lower than Opus 5.5's true standalone capability.
  • AutomationBench was conducted by Zapier and run without model fallbacks; safeguard interventions were treated directly as failures.
  • The standard error for Terminal-Bench 4.0 is reported as ±2.6 points for Opus 5.5.

Another striking remark came from the company itself, noting that "at this level of capability, benchmark differences are becoming less reliable as indicators of real-world differences." Based on internal usage, the perceptible gap between Opus 5.5 and Fable 5.1 is narrower than the scores suggest. Rather than interpreting the table's gaps literally, it is closer to reality to view this as "a model virtually on par with Fable 5.1 has arrived at the Opus price point."

Additionally, CursorBench 4.0 and GDPval-AA v2.1 are different versions from the CursorBench 3.2.0 and GDPval-AA v2 covered in my previous Fable 5.1 article. This accounts for differences in Fable 5.1's figures compared to the previous article, and the figures across the two posts cannot be directly compared.

Strong Even at Lower Effort Levels, Delivering Massive Cost Differences

As the announcement states, "Opus 5.5's clearest advantage lies in efficiency," and this is likely what will have the biggest impact on daily workflows. Thanks to both a lower per-token price and reduced token consumption per task, the combined effect results in an approximate 40% cost reduction compared to Opus 5.

"Effort" is a setting that determines how deeply the model deliberates before responding, offering five tiers: low, medium, high, xhigh, and max. Increasing it leads to deeper reasoning at the expense of tokens and time, whereas lowering it makes responses faster and cheaper. Generally, higher effort yields higher scores. The official graph plots cost per task (log scale) on the horizontal axis against score on the vertical axis, connecting data points for each effort level with a line. Points closer to the top-left indicate "cheaper and higher-scoring."

Terminal-Bench 4.0 scores and costs by effort level. The line for Opus 5.5 is positioned further up and to the left than Fable 5.1, Opus 5, GPT-6 Astra, or GPT-5.6 Sol (Source: Introducing Claude Opus 5.5 \ Anthropic)

In the Terminal-Bench 4.0 graph, the line for Opus 5.5 sits higher and further to the left than all other models. Reading from the scale, Opus 5.5 reaches roughly 57% at medium (around $3 per run) and roughly 64% at high (just under $4). At medium effort, it already outperforms Opus 5 at max (roughly 52%, around $15) and Fable 5.1 at max (55.8%, just under $20) at a fraction of the cost. The official description confirms that Opus 5.5 at default effort outperforms Opus 5 at max effort at roughly one-fifth the cost, while matching GPT-6 Astra at roughly 40% of the cost.

Interestingly, Opus 5.5's score peaks at xhigh (66.4%) and dips slightly at max (roughly 65%). This graph illustrates that "higher effort isn't always strictly better."

GDPval-AA v2.1 Elo and costs by effort level. Opus 5.5 at medium outperforms GPT-6 Astra at max, and xhigh and max are higher than Fable 5.1 at max (Source: Introducing Claude Opus 5.5 \ Anthropic)

GDPval-AA v2.1 evaluates agents on knowledge work across 44 professions, provided by Artificial Analysis. Here as well, the Opus 5.5 curve sits consistently in the top-left. At the official default effort (medium), it scores around 1575 (roughly $0.85 per task), outperforming GPT-6 Astra at max (1542, around $4.50) at about one-fifth the cost. By xhigh (roughly 1820, just over $4), it scores higher than Fable 5.1 at max (1735, around $9.50) at less than half the cost.

These numbers are my estimates read from the chart axes; the underlying measurements were conducted by Anthropic. The announcement also notes that on FrontierCode, Opus 5.5 at default effort outperforms GPT-6 Astra at roughly 20% of the cost, and on CursorBench, it beats GPT-5.6 Sol by 11 points at roughly one-third the cost.

Reports from early testers align with these findings:

  • Deloitte reported that even at its lowest effort level, Opus 5.5 caught 72% of known bugs during code reviews, outperforming Opus 5 at high effort (56%) with fewer false positives and significantly lower output volume.
  • Factory stated this was the first model where they felt comfortable making medium effort the default, matching high-effort Opus 5 results with a 20–25% reduction in output tokens.
  • Optiver noted that on agentic coding tasks, it achieved the same quality as Opus 5 in roughly half the turns, time, and output tokens, reducing costs by 40–50%.

In Claude Code, effort can be adjusted using the arrow keys from the /model screen. It has become much easier to run routine work at medium or high, dialing it up only when tackling difficult tasks.

Pricing: Input/Output Reduced by 20%, Cache Read Slashed by 60%

The pricing is as follows:

Item (per 1M tokens)Opus 5.5Opus 5
Input$4$5
Output$20$25
Cache Read$0.20$0.50
Cache Write$5$6.25

Input and output prices are down 20%, while cache reads (reading cached prompt prefixes) have been discounted by 60%. According to Anthropic, cache reads account for the vast majority of costs in agentic workflows and coding, so this should deliver significant savings for long Claude Code sessions.

For reference, Fable 5.1 is priced at $10 for input, $50 for output, and $0.25 for cache reads. In terms of unit pricing for inputs and outputs, Opus 5.5 is 40% the price of Fable 5.1.

Fast mode is also supported in Claude Code and the Claude Platform. Offering up to 2.5x speed, it is priced at $8 for input and $40 for output (per 1M tokens).

Expanded Usage Limits for Subscriptions

For subscription users, there is good news: alongside the price cuts, 5-hour usage limits are being raised across Pro, Max, Team, and seat-based Enterprise plans.

Additionally, subscribers will receive one Rate Limit reset credit that can be held and used whenever needed. Having an on-demand reset to save for critical deadlines or crunch times is a very welcome touch.

Coding: Excelling at Long-Horizon, Expansive Tasks

The announcement highlights broad, long-horizon jobs—such as codebase-wide migrations and audits—as tasks where Opus 5.5 particularly excels. Examples shared include:

  • An early tester completed an audit and remediation of a 200,000-line codebase in under 3 hours. Opus 5 took over 20 hours and consumed 2.5 times more tokens.
  • In an internal Anthropic test rewriting the HAProxy web server load balancing software from C to Rust, both Opus 5.5 and Fable 5.1 passed virtually all of HAProxy's regression test suite. Opus 5.5 finished in 9.5 hours (compared to 12 hours for Fable 5.1) at 51% lower cost.
  • In a benchmark task optimizing full-page load times across an entire web application, Opus 5.5 succeeded in 39 out of 40 attempts. Opus 5 achieved smaller improvements and broke application behavior along the way.

Here are summaries of feedback from early testers particularly relevant to Claude Code users:

  • GitHub's CPO noted that in tests with GitHub Copilot CLI and VS Code, Opus 5.5 proved to be one of the leanest models measured in terms of tokens and steps, completing more terminal tasks in VS Code in less than half the steps of Opus 5.
  • A Stripe engineer described using a single Opus 5.5 session to coordinate over a dozen child sessions to rebase 40 stacked pull requests over several days, clearly structuring all merge conflicts so that all 40 passed CI by the following afternoon.
  • A Clio engineer reported assigning a massive task spanning six repositories to run unsupervised overnight; the model remained on track for over 18 hours, reached milestones faster than Opus 5, and required almost no human touch-ups.
  • Column noted that the model is significantly better at delegating to sub-agents and proactively devising ways to verify its own work.

While these are individual company reports under differing environments rather than standardized benchmarks, the consistent themes—requiring fewer steps, sustaining long autonomous runs, and delegating well to sub-agents—offer valuable signals for anyone delegating long tasks to Claude Code.

Security Defenses

Running coding agents autonomously inside enterprise infrastructure for extended periods requires verifiable trust. Anthropic positions Opus 5.5 as its "safest coding agent yet," citing three core layers: classifiers that review every operation prior to execution, open-source sandboxes that security teams can audit, and pre-merge code reviews designed to flag vulnerabilities.

The model's intrinsic defenses have also improved. Its resistance to prompt injection (attacks attempting to hijack model behavior via instructions embedded in external data) was on par with or better than Opus 5 across coding, tool use, computer use, and web browsing configurations. In benchmarks by AI security firm Gray Swan, Opus 5.5 tied Fable 5.1 for the lowest prompt injection success rate among tested models.

Knowledge Work: Hallucination-Free Research Stands Out

Among the knowledge work demonstrations, the accuracy of research reports was particularly striking. In an internal test, the model was tasked with writing a quarterly earnings report for a company relying solely on a web copy where the official earnings release was obscured. When evaluated automatically with a strict zero-tolerance criteria for fabricated numbers or citations, Opus 5.5 passed in 16 out of 18 attempts across varied effort levels, whereas Fable 5.1 and Opus 5 scored zero passes.

Other examples included:

  • Analyzing a proposed merger between two hypothetical HR software vendors, building an Excel financial model, and generating an executive presentation deck. Opus 5.5 finished in 63 minutes (vs. 93 minutes for Opus 5) at 50% lower cost. While their conclusions aligned, Opus 5.5 produced more meticulous modeling and clearer slides.
  • Investment firm Walleye Capital reported that Opus 5.5 solved evaluation tasks even at lowest effort, and at high effort, it detected and corrected a one-minute index offset error embedded in the prompt instructions—something no prior model had caught.
  • Data analytics platform Hex presented a task to diagnose whether a shipment issue was a package delay or tracking telemetry failure; where Opus 5 reviewed the delivery logs and reported "tracking nominal," Opus 5.5 accurately determined that both the shipment was delayed and the tracking system was broken.

Rather than stopping at the first plausible answer, it digs deeper. I noted in my Opus 5 write-up that the model's "self-verification capability" had improved; that trajectory appears to have accelerated further.

Cleaner, More Readable Writing

A frequent complaint regarding Opus 5 was the readability of its prose. Opus 5.5 addresses this directly.

According to Anthropic, Opus 5.5 produces messages that are scannable at a glance, front-loads vital conclusions, cuts down on technical jargon and stylized phrasing, and adheres more strictly to user-provided formatting guidelines. On the release page, sample outputs for explaining a bug fix are compared side by side:

  • Opus 5 walks sequentially through its diagnostic steps—"what was found," "bug details"—meaning the reader only grasps the severity at the very end.
  • Opus 5.5 leads with a clear headline: "Additional decline was caused by billing refactoring bug," presents the dollar impact upfront ($1.50 from tier changes, $9.92 from the bug), and only then transitions to the code diff.

An engineer at early tester Ramp noted that verbose, hard-to-follow outputs were their biggest frustration with frontier models, and Opus 5.5 resolves that issue. Box reported a 40% reduction in verbosity while consuming one-third the tokens compared to Opus 5.

Anthropic highlights that readability is not just an ergonomics upgrade, but a safety feature: cleaner prose makes the agent's actions easier to monitor and verify. Spending less time deciphering long Claude Code logs might end up being the most noticeable day-to-day improvement.

Safety and Alignment

The safety section opens with the philosophy of "Pacing the frontier." This follows Dario Amodei's statement the prior week advocating that the pace of AI advancement should be actively paced so that safety mechanisms stay ahead of capabilities. Anthropic outlines a two-horizon framework: refining mitigations for current models while establishing stricter safety standards and public policy coordination for future models capable of automating AI research itself.

Key takeaways from Opus 5.5's alignment evaluations include:

  • Automated Behavioral Audits - In the primary evaluation spanning roughly 2,000 scenarios, it outperformed recent Claude models across nearly all metrics. It also ranked as the company's strongest model across several honesty benchmarks.
  • Boundary-Transgressing Tendencies - In a new evaluation measuring propensity to exceed designated containment boundaries, the frequency of attempts to circumvent boundaries fell roughly 85% compared to Opus 5 and Mythos 5.1; all observed attempts were low-severity and self-reported by the model.
  • Mitigation of Incident Behaviors - Notable reductions were observed in behaviors linked to recent cybersecurity incidents, including motivated reasoning, sandbox escape attempts, and malicious actions taken after concluding "this is a simulation."

The release also candidly shares limitations: Opus 5.5 frequently deduces when it is undergoing evaluation, making it harder to predict real-world behavior purely through benchmark testing. Anthropic notes that crafting evals that catch every failure mode prior to deployment remains an unsolved challenge. When running Claude Code autonomously with broad permissions, it remains prudent to define explicit operational boundaries rather than relying solely on safety improvements.

Same Safeguard Tier as Fable 5.1

Opus 5.5 is the first Opus-class model to deploy with the same tier of safeguards for cybersecurity, biology, and anti-distillation as Fable 5.1. The announcement clarifies that when these safeguards trigger, the model switches transparently to a fallback model:

  • Cybersecurity - Given its strong cyber capabilities, Fable 5.1-equivalent safeguards are in place. While standard software debugging and bug fixes operate normally, most dedicated cybersecurity tasks are routed to Opus 4.8. For defensive specialists, Opus 5.5 will soon be added to the Cyber Verification Program, creating a three-tiered access structure that includes Mythos models.
  • Biology - Outperforming Opus 5 across many domains and matching or exceeding Mythos 5.1, it incorporates the same biological safeguards as Fable 5.1. Organizations conducting legitimate R&D can apply via the Life Sciences Verification Program.
  • Anti-Distillation - The preserved thinking mechanism introduced in Fable 5.1 is now enabled on Opus 5.5. Designed to prevent API users from manipulating past context to extract raw internal reasoning traces, it applies to API accounts created on or after August 31, 2026. Developers building integrations with custom conversation history pipelines should consult the Help Center article.

For Claude Code users, the main practical implication is that certain security-adjacent tasks may silently fall back to Opus 4.8. Since this is the first time an Opus model features Fable-tier safeguards, anyone previously running security investigations on Opus 5 should keep this behavioral change in mind.

Data Retention and Availability

As with prior Opus models, Opus 5.5 is available under zero data retention (ZDR) terms.

Two other operational changes of note:

  • Like Fable 5.1, output includes watermarking to comply with the EU AI Act.
  • Running the model with thinking disabled is no longer supported.

Availability began on the day of announcement across all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. The API model ID is claude-opus-5-5.

Trying It Out in Claude Code

In my setup, opening /model in Claude Code showed that the Default (recommended) option had already been updated to Opus 5.5.

Claude Code /model screen showing Default (recommended) and Opus (1M context) both set to Opus 5.5 with 1M context

The first entry, Default, and the second, Opus (1M context), both map to "Opus 5.5 with 1M context," labeled for "everyday to complex tasks." The third option is Fable 5.1, followed below by Sonnet 5 and Haiku 4.5. Opus 5 no longer appears in the list; using the older model now requires passing the model ID explicitly via --model at startup. Effort can also be specified at launch via --effort.

claude --model claude-opus-5
claude --effort medium

If the update hasn't reflected locally yet, try updating and restarting the CLI:

claude update

Model visibility and defaults may vary depending on your plan, provider, and administrator configurations. If you are still running Claude Code via Homebrew or npm, switching to the native installer ensures smooth automatic updates as new models land. A walkthrough is available in How Switching Claude Code from Homebrew to Native Install Improved My Workflow.

Opus 5 vs. Opus 5.5 Comparison

Here is a summary of the core differences:

ItemOpus 5Opus 5.5
Pricing (Input/Output per 1M tokens)$5 / $25$4 / $20
Cache Read (per 1M tokens)$0.50$0.20
Typical Workload CostBaseline~40% reduction (default settings)
Output Generation SpeedBaseline>30% faster
Terminal-Bench 4.052.3%66.4%
CursorBench 4.046.6%57.8%
GDPval-AA v2.117081846
Cyber / Bio / Anti-Distillation SafeguardsNot Fable 5.1 tierFable 5.1 equivalent (falls back to Opus 4.8, etc., upon intervention)
API Model IDclaude-opus-5claude-opus-5-5

Benchmark figures are based on the comparison tables in the official announcement.

Why This Matters for Users

Reframing the above points through the lens of daily development work:

1. Fable 5.1-class performance at Opus pricing

It delivers parity with Fable 5.1 across most tasks and beats it on Terminal-Bench 4.0, all at 40% of Fable 5.1's token pricing. Workflows previously split into "Fable for hard tasks, Opus for routine work" can largely consolidate onto Opus 5.5.

2. Cheaper, faster execution for identical tasks

The combination of reduced token rates and higher token efficiency lowers typical workload costs by roughly 40%. Generation speeds are over 30% faster, and subscription usage limits have been expanded.

3. Dramatically more readable summaries

While subtle, this makes a huge difference in day-to-day Claude Code usage. Having the conclusions stated upfront materially reduces the cognitive overhead of reviewing long agent sessions.

4. Behavioral differences on security-related tasks

A caveat to keep in mind: because Opus 5.5 incorporates Fable 5.1-tier safeguards, parts of security-sensitive workflows will be redirected to Opus 4.8. Treating it as an exact drop-in for Opus 5 in security domains may lead to unexpected results.

Summary

The key takeaways for Claude Opus 5.5 are:

  • First model in the Claude 5.5 family. Delivers performance on par with Fable 5.1 across many tasks. Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks.
  • Terminal-Bench 4.0 reached 66.4% (vs. 55.8% for Fable 5.1, 52.3% for Opus 5), CursorBench 4.0 hit 57.8%, and GDPval-AA v2.1 reached 1846.
  • GPT-6 Astra retains the lead on AutomationBench and Terminal-Bench-Science 0.1.
  • Default effort on Opus 5.5 outperforms max effort Opus 5 at roughly one-fifth the cost (Terminal-Bench 4.0).
  • Pricing is $4 input, $20 output, and $0.20 cache read (per 1M tokens). Runs ~40% cheaper than Opus 5 on typical workloads, with over 30% faster generation.
  • Fast mode provides up to 2.5x speed at $8 input / $40 output.
  • Raised 5-hour usage limits for Pro, Max, Team, and seat-based Enterprise plans, plus one on-demand Rate Limit reset credit.
  • Output text is more readable, front-loading critical conclusions.
  • Highest scores to date on automated behavioral audits; boundary avoidance attempts decreased ~85% compared to Opus 5.
  • Includes Fable 5.1-class safeguards for cyber, bio, and anti-distillation; many cyber tasks route to Opus 4.8.
  • Available with zero data retention. Thinking cannot be disabled.
  • API model ID is claude-opus-5-5. Available day-one across AWS, Google Cloud, and Azure.

It has been only three weeks since my Fable 5.1 post noted that "Fable pulled ahead again after Opus 5 briefly caught up." Now, Opus has already caught up to Fable 5.1's capabilities at a lower price point. The pattern of pioneering breakthroughs in higher tiers and cascading them down to affordable price points within weeks has become unmistakable.

I plan to make Opus 5.5 my daily driver in Claude Code, alternating between medium and high effort. I'll share further observations in a future post once I have more mileage on it.

Reference Links

Share this article

Related Articles