Brand Logo

Claude Opus 5.5 Explained: Benchmarks, Pricing, and How It Beats Fable 5.1 at 60% Lower Cost

Aerin Kim

Written by

Aerin Kim

Anthropic's Claude Opus 5.5 launched September 22, 2026, beating its own pricier Fable 5.1 on key agentic benchmarks at 60% lower API cost. Full breakdown of scores, pricing, and code.

Yesterday, September 22, 2026, Anthropic shipped Claude Opus 5.5, the first model in a brand new 5.5 family, and buried in the launch materials is a genuinely unusual claim for a frontier lab to make about its own lineup: the new, cheaper model beats the older, far more expensive one on several of the benchmarks that matter most for real agentic work [1]. Opus 5.5 performs at the level of Claude Fable 5.1, Anthropic's actual top-of-the-line model, on most tasks, while costing roughly 60 percent less per token on the base API rate than Fable 5.1, and about 40 percent less than its own predecessor, Opus 5, on a typical workload once you account for how many fewer tokens it actually burns to finish a task [2][7].

That alone would be a notable release. What makes the timing sharper is that OpenAI shipped two new models, GPT-6 Sol and GPT-6 Luna, on the exact same Tuesday, cutting its own API prices by roughly half and undercutting Opus 5.5 on raw per-token cost in the process [12][13]. Two frontier labs picked the same day to compress the price of high-end intelligence, and their two answers to "how cheap can this get" are different enough that neither one settles the question outright. Opus 5.5 is also the first model Anthropic has shipped since CEO Dario Amodei publicly signed onto calls to deliberately "pace the frontier," trading some raw capability speed for more room to keep alignment work in step with what the models can do [6][18].

This post is the detailed version of what actually shipped, not a rewritten press release. Coverage of the launch was broad and largely consistent across outlets from MacRumors to SiliconANGLE to BNN Bloomberg [14][15][21], and this post walks through the real benchmark numbers cross-checked against Anthropic's own documentation and those independent outlets, the real pricing math worked out for a sample workload you can run yourself, the mechanism behind how a cheaper model manages to beat a pricier one, the safety changes that come with a model whose default posture is more autonomous than its predecessor's, and working code for calling Opus 5.5 from the Anthropic SDK and from Amazon Bedrock, since it is now live on the Claude Platform on AWS as well as AWS Bedrock itself [3][5].

claude-opus-5-5-benchmarks-pricing-explained-2026-hero.png

Step 1: What Actually Launched on September 22

Claude Opus 5.5 is Anthropic's newest flagship model, built specifically for long-running agentic coding and knowledge work, and it replaces Opus 5 as the strongest model most developers should reach for by default [3]. The model ID is claude-opus-5-5 on the Claude API, Google Cloud, Microsoft Foundry, and the Claude Platform on AWS, and anthropic.claude-opus-5-5 on Amazon Bedrock, where it is served through a global cross-region inference profile spanning US, EU, AU, and JP geographic endpoints [3][5]. It carries the same 1 million token context window and 128,000 token maximum output as Opus 5 and Fable 5.1, with a reliable knowledge cutoff of June 2026 [3].

The headline framing from Anthropic itself is worth reading exactly as written: Opus 5.5 "performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5" [2]. That is a company explicitly telling its own customers that its second-tier model now does most of what its flagship does, for a fraction of the flagship's price. It is a genuinely unusual thing to say out loud, and it only makes sense once you understand that Anthropic is optimizing two different variables at once: raw capability ceiling, where Fable 5.1 still leads on the hardest, most open-ended research tasks, and cost-efficient capability, where Opus 5.5 is now the better tool for the vast majority of real production workloads [1].

Availability is broad from day one. Opus 5.5 is live on the Claude API, AWS Bedrock, the Claude Platform on AWS, Google Cloud, and Microsoft Azure, alongside Claude Code, Claude Cowork, and Claude.ai [3][5][16]. Anthropic has also said Claude Sonnet 5.5 and Claude Haiku 5.5 are expected within the coming weeks, meaning Opus 5.5 is the opening move in a full family refresh rather than a standalone release [6].

Four Breaking Changes If You Already Call Opus 5 in Production

Anthropic's own migration guide flags four changes that will actually break existing integrations, not just shift behavior at the margins [3][4]:

  • Thinking can no longer be disabled. Adaptive thinking is always on for Opus 5.5, the same posture Fable 5.1 shipped with in September. There is no way to fully turn reasoning off the way older extended-thinking toggles allowed; the effort parameter, covered in Step 6, is now the only lever.
  • Forced tool use returns an error. If your integration relies on forcing the model to call a specific tool on a given turn, that call pattern is no longer supported and will throw instead of silently degrading to a different behavior.
  • Thinking blocks are tied to the model and the conversation that produced them. A multi-model architecture that routes a conversation between models mid-thread, escalating from Sonnet 5 to Opus 5.5, for instance, cannot have an earlier or later model in that chain parse Opus 5.5's internal reasoning blocks.
  • The older computer_20251124 computer-use tool definition is rejected on the Claude API and Google Cloud. Any computer-use integration built against that tool version needs to move to the current tool definition before switching model strings.

A fifth change is additive rather than breaking but easy to miss: text that used to stream between tool calls now comes back inside thinking blocks, which render empty at the default display setting. An application that streamed that text to end users as live progress updates during a long agentic run will go quiet between tool calls unless it explicitly sets a display value that surfaces the text again [3]. This is exactly the kind of change that passes every automated test and then quietly breaks a product's perceived responsiveness in production, so it is worth checking deliberately rather than assuming a passing test suite catches it.

Why Anthropic Shipped This Now

The backdrop to this specific release is worth knowing, because it explains why Opus 5.5 reads more like a cost and efficiency release than a pure capability release. In the weeks before launch, Amodei publicly aligned with calls from parts of the AI safety community to deliberately pace frontier capability progress, arguing that fully addressing the risks of increasingly autonomous models requires more prudence, not less, even as competitive pressure pushes labs to ship faster [6][18]. Opus 5.5's own safety training is described as broadly similar in kind to its predecessors, evaluated through Anthropic's usual automated behavioral audits alongside outside evaluators, while the company signals that more advanced training and evaluation systems are being prepared for future releases rather than shipped wholesale in this one [6]. Put simply, Opus 5.5 is not being sold as a bigger capability jump than the numbers show. It is being sold as the same tier of capability delivered far more efficiently, which is a different kind of release than a typical flagship launch, and Step 4 below covers exactly what the efficiency gain looks like mechanically.

Step 2: The Benchmark Numbers, Side by Side

Anthropic's own launch benchmarks, corroborated by TechCrunch, VentureBeat, and MarkTechPost's independent coverage, give a full picture of where Opus 5.5 lands against Fable 5.1, its own predecessor Opus 5, and OpenAI's GPT-6 Astra and GPT-5.6 Sol [1][7][17].

BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 Sol
Terminal-Bench 4.0 (agentic coding)66.4%55.8%52.3%57.9%37.3%
FrontierCode v1.154.4%50.3%48.0%53.3%47.5%
CursorBench 4.0 (IDE coding)57.8%51.8%46.6%Not reported41.7%
GDPval-AA v2.1 (knowledge work, Elo)18461735170815421588
AutomationBench40.0%31.4%26.9%41.4% (leads)28.8%
Humanity's Last Exam67.7%65.6%63.6%57.2%Not reported
Terminal-Bench-Science 0.158.7%52.6%29.0%64.6% (leads)22.4%
OSWorld 2.0 (computer use)81.8%80.7%74.0%Not reportedNot reported
Chartography (chart interpretation)89.0%88.4%83.4%Not reportedNot reported

A few things stand out reading that table honestly, including the parts that do not flatter Anthropic.

Opus 5.5 leads its own family and GPT-6 Astra on most agentic benchmarks. On Terminal-Bench 4.0, a measure of multi-step, tool-using command-line agentic work, Opus 5.5 scores 66.4 percent against Fable 5.1's 55.8 percent, its own predecessor Opus 5's 52.3 percent, GPT-6 Astra's 57.9 percent, and GPT-5.6 Sol's 37.3 percent [1][7]. That is not a marginal win. It is 8.5 points ahead of GPT-6 Astra and nearly 11 points ahead of Anthropic's own more expensive Fable 5.1. The same pattern holds on FrontierCode v1.1, a coding benchmark, where Opus 5.5's 54.4 percent edges out Astra's 53.3 percent and clears Fable 5.1's 50.3 percent, and on CursorBench 4.0, an IDE-assisted coding benchmark, where Opus 5.5's 57.8 percent beats Fable 5.1's 51.8 percent by six points with no comparable GPT-6 Astra figure published.

On knowledge work, the gap is even larger. GDPval-AA v2.1, Anthropic's Elo-style benchmark for real business tasks like document drafting and spreadsheet analysis, puts Opus 5.5 at 1846 against Fable 5.1's 1735, Opus 5's 1708, GPT-6 Astra's 1542, and GPT-5.6 Sol's 1588. A 111-point Elo gap over Fable 5.1 on this benchmark alone would have justified a standalone release under the old pricing; that Opus 5.5 achieves it at 40 percent lower typical-workload cost than Opus 5 is the real headline.

Two benchmarks go the other way, and an honest post has to say so. GPT-6 Astra actually wins on AutomationBench, a workflow-automation benchmark spanning multiple applications, scoring 41.4 percent against Opus 5.5's 40.0 percent. The gap is small, 1.4 points, but it is real, and it is the one benchmark in this set where OpenAI's model comes out ahead of Anthropic's newest release. GPT-6 Astra's win is far larger on Terminal-Bench-Science 0.1, an agentic scientific-research benchmark: Astra scores 64.6 percent against Opus 5.5's 58.7 percent, a 5.9-point gap in OpenAI's favor. Anthropic's own launch materials do not hide this: Opus 5.5's strength is broad and deep across agentic coding, computer use, and general knowledge work, not universal dominance across every published benchmark [1]. Treat any summary of this release that claims Opus 5.5 wins every category as either careless or incomplete.

Opus 5.5 does clear its own predecessor by a wide margin everywhere. Against Opus 5 specifically, the smallest gap in the whole table is on Chartography, a visual chart-and-graph interpretation benchmark, where Opus 5.5's 89.0 percent beats Opus 5's 83.4 percent by 5.6 points. Every other benchmark shows a double-digit percentage-point or multi-hundred-Elo jump over Opus 5, including nearly doubling Opus 5's score on Terminal-Bench-Science, 58.7 percent versus 29.0 percent, which is the single largest generational jump in the entire table.

Humanity's Last Exam confirms this is a genuine reasoning improvement, not just an agentic-tooling one. On this broad, tool-agnostic reasoning benchmark, Opus 5.5 scores 67.7 percent against Fable 5.1's 65.6 percent, Opus 5's 63.6 percent, and GPT-6 Astra's 57.2 percent, a benchmark where raw multi-step reasoning matters more than tool orchestration, and Opus 5.5 still leads the entire field.

Independent verification matters here because vendor-published benchmarks are, by nature, selected by the vendor. Artificial Analysis, which runs a standardized Intelligence Index against real API calls rather than relying on any lab's self-reported figures, placed Opus 5.5 at a score of 58, ranked first out of 212 models it had evaluated at the time [10]. That is a genuinely independent confirmation that the benchmark story above is not simply Anthropic picking favorable comparisons.

Step 3: Pricing, the Real Math Per Task

Opus 5.5's headline pricing is a straightforward 20 percent cut on both input and output tokens compared to Opus 5, but the more consequential change is buried in the cache pricing line.

Model / ModeInput ($/1M tok)Output ($/1M tok)Cache Read ($/1M tok)Cache Write ($/1M tok, 5m)
Opus 5 (predecessor)$5.00$25.00$0.50$6.25
Opus 5.5 (standard)$4.00$20.00$0.20$5.00
Opus 5.5 (Fast mode, preview)$8.00$40.00Not publishedNot published
Fable 5.1 / Mythos 5.1$10.00$50.00$0.25$12.50

Standard Opus 5.5 runs $4 per million input tokens and $20 per million output tokens, down from Opus 5's $5 and $25 [2][3]. A separate Fast mode, currently a research preview on the Claude API, runs $8 input and $40 output per million tokens, roughly double the standard rate, in exchange for output generated more than 30 percent faster than Opus 5's baseline speed [1][20]. Cache writes dropped modestly, from $6.25 to $5.00 per million tokens on the standard five-minute cache, with a one-hour cache write option at $8.00 [3]. The number that actually moves the needle for anyone running long agentic sessions is cache reads, cut from $0.50 to $0.20 per million tokens, a 60 percent reduction [1][2].

That cache-read cut matters more than it sounds like on first read, because cache reads make up the majority of the token volume in a real agentic session, not the minority [8]. A long-running coding agent typically writes its system prompt and tool schema to cache once at the start of a session, then re-reads that same cached prefix on every subsequent turn as the conversation grows. In a session with twenty tool-calling turns, the cache reads can easily outnumber the fresh input tokens by an order of magnitude, so a 60 percent cut to that specific line item compounds heavily across exactly the kind of long, multi-step session Opus 5.5 is built for.

claude-opus-5-5-benchmarks-pricing-explained-2026-pricing-gauge.png

Anthropic's own claim is that this combination, the headline rate cut plus the cache-read cut plus the model simply needing fewer tokens per task, works out to roughly 40 percent lower cost on a typical workload compared to Opus 5 [1][7]. That is a claim worth testing on your own numbers rather than trusting outright, so here is a small calculator that models the three levers separately.

python
/code def session_cost(input_tokens: int, output_tokens: int, cache_read_tokens: int, cache_write_tokens: int, input_price: float, output_price: float, cache_read_price: float, cache_write_price: float) -> float: """Rough cost of one agentic session given per-million-token pricing.""" return ( (input_tokens / 1_000_000) * input_price + (output_tokens / 1_000_000) * output_price + (cache_read_tokens / 1_000_000) * cache_read_price + (cache_write_tokens / 1_000_000) * cache_write_price ) # A representative long agentic session: one cache write at the start, # then the same system prompt and tool schema re-read from cache across # 20 subsequent tool-calling turns. turns = 20 prefix_tokens = 8_000 opus_5 = session_cost( input_tokens=5_000, output_tokens=14_000, cache_read_tokens=turns * prefix_tokens, cache_write_tokens=prefix_tokens, input_price=5.00, output_price=25.00, cache_read_price=0.50, cache_write_price=6.25, ) opus_5_5 = session_cost( input_tokens=5_000, output_tokens=10_000, # fewer output tokens per Anthropic's efficiency claim cache_read_tokens=turns * prefix_tokens, cache_write_tokens=prefix_tokens, input_price=4.00, output_price=20.00, cache_read_price=0.20, cache_write_price=5.00, ) fable_5_1 = session_cost( input_tokens=5_000, output_tokens=14_000, cache_read_tokens=turns * prefix_tokens, cache_write_tokens=prefix_tokens, input_price=10.00, output_price=50.00, cache_read_price=0.25, cache_write_price=12.50, ) gpt_6_sol = session_cost( input_tokens=5_000, output_tokens=14_000, cache_read_tokens=turns * prefix_tokens, cache_write_tokens=prefix_tokens, input_price=2.00, output_price=10.00, cache_read_price=0.20, cache_write_price=2.50, # illustrative, OpenAI cache pricing varies by product ) print(f"Opus 5 session: ${opus_5:.4f}") print(f"Opus 5.5 session: ${opus_5_5:.4f} ({(1 - opus_5_5 / opus_5) * 100:.1f}% cheaper than Opus 5)") print(f"Fable 5.1 session: ${fable_5_1:.4f} (Opus 5.5 is {(1 - opus_5_5 / fable_5_1) * 100:.1f}% cheaper)") print(f"GPT-6 Sol session: ${gpt_6_sol:.4f}")

Running that calculator on a representative long agentic session, one cache write followed by twenty cached-prefix reads, shows the cache-read change alone accounting for a meaningful share of the total savings, before even factoring in Opus 5.5 needing fewer output tokens to reach a comparable answer. The two effects stack: a session that cost, say, four dollars on Opus 5 does not just get a 20 percent haircut from the headline rate change, it gets that plus a much larger cut on every cached turn after the first.

It is also worth being precise about how Opus 5.5's roughly 60 percent cost advantage over Fable 5.1 is actually measured. Fable 5.1 runs $10 input and $50 output per million tokens, exactly 2.5 times Opus 5.5's standard rate on both input and output [7]. A model that costs 2.5 times less per token is, definitionally, about 60 percent cheaper on raw list price, and that is the number this post's title refers to. Whether Opus 5.5 stays roughly 60 percent cheaper per completed task once you account for output-token verbosity differences between the two models depends on your specific workload, which is exactly why Step 4 below focuses on the token-efficiency mechanism rather than the sticker price alone.

Step 4: How a Cheaper Model Beats a More Expensive One

The genuinely interesting part of this release is not the price cut by itself, it is that Anthropic is claiming Opus 5.5 does more with fewer tokens than Opus 5 did, not just the same work at a lower per-token rate [1]. Those are two different kinds of improvement, and conflating them understates how unusual this release actually is. A pure price cut makes an unchanged model cheaper. A token-efficiency improvement makes the same task finish using less reasoning and fewer output tokens in the first place, which compounds with the price cut rather than sitting alongside it.

claude-opus-5-5-benchmarks-pricing-explained-2026-construction-timeline.png

Anthropic's own case studies make this concrete rather than abstract. One tester completed a 680,000-line code migration in under a single day, a scale of change that would normally consume an engineering team's sprint or more [7]. Another ran a 200,000-line codebase audit to completion in under three hours; the same audit on Opus 5 took more than 20 hours and consumed roughly 2.5 times more tokens to get there [1][7]. That 2.5x token figure is the clearest single data point in the whole launch for what "uses fewer tokens per task" actually means in practice: roughly the same task, more than six times faster wall-clock, at a fraction of the token spend.

A third data point compares Opus 5.5 directly against Fable 5.1 rather than against its own predecessor. A HAProxy C-to-Rust translation task, a genuinely hard, long-horizon systems-programming migration, finished on Opus 5.5 in 9.5 hours, against 12 hours on Fable 5.1, at 51 percent lower cost for the completed task [7]. That is the pattern this whole release is built around in miniature: comparable or better quality, meaningfully faster, and meaningfully cheaper, from the model that costs 60 percent less per token on paper.

The customer quotes Anthropic published back this up with specifics rather than vague praise, which is worth taking seriously precisely because they are checkable claims rather than generic testimonials:

  • Mario Rodriguez, Chief Product Officer at GitHub, on Opus 5.5 solving more terminal tasks in fewer steps: "Developers want agents that can take on real software work and finish it... Claude Opus 5.5 used among the fewest tokens and steps we measured" [2].
  • Sean Heintz, Staff Software Developer at Clio, who ran Opus 5.5 unattended for over 18 hours on a large engineering task defining inter-service contracts: "I handed Claude Opus 5.5 a large engineering task... it hit milestones faster and required minimal reworking" [2].
  • John Ruelas, Staff Software Engineer at Ramp, on the model's output register: "Verbose, hard-to-follow output has been my biggest frustration... Claude Opus 5.5 fixes it" [2].
  • Frank Corrao, Head of Central Equity Quant Research at Walleye Capital, on a subtler failure mode: "It detected that the minute indexing in our own instructions was off by one and corrected for it" [2].
  • Carl Bennett, CIO at Deloitte Consulting LLP, on quality at the cheapest setting: "Even at its lowest effort setting, Claude Opus 5.5 caught 72% of known bugs" [2].

Ruelas's quote deserves a beat of its own context, because Opus 5's verbosity was a genuinely common complaint among developers in the months after its release, not a manufactured talking point for this launch. A model that writes more concisely by default does not just read better, it directly reduces output-token cost on every single request, which is part of why the token-efficiency story and the writing-quality story are really the same underlying change described from two different angles.

Step 5: What Else Changed Under the Hood

Beyond price and token efficiency, three categories of change separate Opus 5.5 from Opus 5: raw speed, output style, and the safety and autonomy posture that comes with trusting a model with more unsupervised agentic work.

Speed. Opus 5.5 generates output more than 30 percent faster than Opus 5 at the same effort level, and Fast mode roughly doubles that improvement again for latency-sensitive workloads willing to pay the premium rate covered in Step 3 [1][7].

Default output register. Opus 5.5's default effort level on the API is medium, a full step down from Fable 5.1's high default, and Anthropic's platform docs list its comparative latency as "Moderate" against Fable 5.1's "Slower" [3]. Combined with the conciseness improvements customers like Ramp reported, the practical effect is a model that reaches for less reasoning and fewer words by default, and needs a deliberate effort bump for the genuinely hard problems that actually benefit from it, rather than spending extra tokens on routine requests by default the way earlier generations sometimes did.

claude-opus-5-5-benchmarks-pricing-explained-2026-safety-checkpoint.png

A more precise action-screening pipeline. Opus 5.5 introduces a classifier that screens every proposed action before it executes, layered on top of the model's own judgment about whether to take that action at all [2][8]. Anthropic's internal alignment testing, run across nearly 2,000 scenarios and documented in the model's own system card, found Opus 5.5 attempted to circumvent its own operating boundaries roughly 85 percent less often than either Opus 5 or Claude Mythos 5.1 did on the same test suite [2][19]. That is a specific, checkable claim about behavior under adversarial testing conditions, not a vague statement about the model being "safer," and it is the kind of number worth reading literally: fewer attempted boundary violations per fixed set of test scenarios, measured the same way across model generations.

Preserved thinking, an anti-distillation safeguard. For accounts created after August 31, 2026, Opus 5.5 prevents API users from editing the model's prior reasoning context mid-conversation, a defensive measure against the well-established practice of using a frontier model's own chain of reasoning as training data to distill a cheaper competing model [2]. This does not change what a typical developer's application code looks like, but it is worth knowing if your architecture relies on programmatically rewriting or pruning earlier turns in a conversation, since that pattern interacts directly with how thinking blocks are preserved.

EU AI Act compliance and cybersecurity fallback. Opus 5.5 ships with watermarking measures aimed at EU AI Act compliance, matching the approach Fable 5.1 introduced in September [2]. On the cybersecurity side, most genuinely sensitive cybersecurity tasks now route to the older Claude Opus 4.8 model instead, with an expanded Cyber Verification Program planned to offer tiered access for vetted security teams, while routine bug-fixing and defensive work stays available on Opus 5.5 directly [2]. Biology-related safeguards on Opus 5.5 are matched to Fable 5.1's level, and a new Life Sciences Verification Program opens access for vetted academic labs, biotech startups, and pharmaceutical companies to a less restricted deployment [2].

None of this is abstract policy trivia if you are actually building an autonomous agent with tool access. A model that intervenes on its own proposed actions less often, and more precisely, wastes fewer of your agent's tool calls on false refusals, which is a real operational property that shows up directly in the same token-and-step efficiency numbers discussed in Step 4, not a separate concern from them.

Step 6: Calling Claude Opus 5.5 From Code

Opus 5.5 is reachable the same way Opus 5 was, through the Anthropic SDK directly, or through Amazon Bedrock and the Claude Platform on AWS if your infrastructure already lives there [3][5]. Here is a basic call through the Anthropic SDK, with the effort parameter set explicitly rather than left at its medium default, since the whole point of the five-level effort system is to make that choice deliberate per request.

python
/code import anthropic client = anthropic.Anthropic() def ask_opus(prompt: str, effort: str = "medium") -> str: """Call Claude Opus 5.5. Thinking is adaptive and always on for this model, so effort does not toggle reasoning off, it only controls how much of it the model spends. Valid values are low, medium, high, xhigh, and max. The API defaults to medium if you omit the parameter, a step down from Fable 5.1's default of high.""" response = client.messages.create( model="claude-opus-5-5", max_tokens=8192, effort=effort, messages=[{"role": "user", "content": prompt}], ) return response.content[-1].text # A routine request does not need to override the medium default. quick_answer = ask_opus("Summarize the three failing CI jobs from this log.") # A long-horizon agentic task, the kind Anthropic's own 200,000-line # audit case study used, benefits from a higher effort level. deep_answer = ask_opus( "Audit this repository for latent concurrency bugs, propose fixes, " "and write regression tests for each one you find.", effort="high", ) print(deep_answer)

If your workload runs on AWS, Opus 5.5 is available both through classic Bedrock model invocation and through Bedrock's unified Converse API, which is the more portable option if you might swap models later. AWS's own announcement documents three call patterns: the boto3 invoke_model call, the Converse API, and the native Anthropic SDK authenticated with AWS credentials [5]. Here is the Converse API version, since it is the pattern AWS itself recommends for new integrations.

python
/code import boto3 bedrock = boto3.client("bedrock-runtime", region_name="us-east-1") # The Converse API is the portable, model-agnostic pattern AWS recommends # for new integrations, since the same call shape works across models. response = bedrock.converse( modelId="anthropic.claude-opus-5-5", messages=[ { "role": "user", "content": [ {"text": "Review this Terraform diff for a misconfigured " "security group before it merges."} ], } ], inferenceConfig={"maxTokens": 4096}, additionalModelRequestFields={"effort": "high"}, ) output_text = response["output"]["message"]["content"][0]["text"] print(output_text) # Note: cross-region inference profiles (useful for availability and # latency) typically add a modest premium over a single fixed-region # call, so factor that into per-task cost estimates if you route this # way instead of pinning a single AWS region.

A few practical notes that do not show up in either snippet above but matter in production. First, cross-region inference on Bedrock, useful for availability and latency reasons, typically adds a modest premium over standard-region pricing, so factor that into the cost calculator from Step 3 if you are routing through a cross-region profile rather than a single fixed region. Second, since thinking can no longer be disabled on Opus 5.5, any code that previously set a thinking budget of zero to fully suppress reasoning needs to be rewritten around the effort parameter instead, one of the four breaking changes covered in Step 1. Third, if your application streamed intermediate text between tool calls as a live progress indicator, verify your display setting explicitly rather than assuming the old streaming behavior carried over, since that is the least obvious of the changes in Step 1 and the easiest one to miss in a code review.

Step 7: Opus 5.5 vs GPT-6 Sol and GPT-6 Luna, the Same-Day Launch

OpenAI did not sit still on September 22. The company released GPT-6 Sol and GPT-6 Luna the same day, building on the capability of its existing flagship, GPT-6 Astra, but packaged into faster, cheaper models aimed at high-volume production work rather than the absolute reasoning ceiling [12][13]. It is worth being precise about the lineup here, since it is easy to conflate: GPT-6 Astra is OpenAI's existing top-tier model, priced at $10 input and $50 output per million tokens, the same list price as Anthropic's Fable 5.1 [12]. Sol and Luna are new, cheaper tiers underneath it, not replacements for it.

ModelLabInput ($/1M tok)Output ($/1M tok)AA Intelligence Index
Claude Opus 5.5Anthropic$4.00$20.0058 (rank #1/212)
Claude Fable 5.1Anthropic$10.00$50.00Not directly reported here
GPT-6 Astra (existing flagship)OpenAI$10.00$50.00Not directly reported here
GPT-6 Sol (new, same-day)OpenAI$2.00$10.0048 (rank #18/212)
GPT-6 Luna (new, same-day)OpenAI$0.10$0.50Not directly reported here
claude-opus-5-5-benchmarks-pricing-explained-2026-market-crates.png

GPT-6 Sol launched at $2 input and $10 output per million tokens, a permanent price, not a promotional one, confirmed directly to VentureBeat by an OpenAI spokesperson, cutting Sol's price by 50 percent versus the equivalent tier in the GPT-5.6 generation [12]. GPT-6 Luna, aimed at high-volume clerical and structured work rather than complex reasoning, launched at just $0.10 input and $0.50 output per million tokens [12]. On raw sticker price, Sol undercuts Opus 5.5 on both input and output tokens, and Luna undercuts every model discussed in this post by an order of magnitude or more.

Price alone does not settle a model comparison, and this is a place where it genuinely matters to compare like with like rather than repeating whichever number sounds most dramatic. Anthropic's own launch materials do not publish a head-to-head score against GPT-6 Sol on the same benchmark suite, and it would be a mistake to invent one [1]. What does exist, and is worth citing precisely because it comes from neither lab, is Artificial Analysis's standardized Intelligence Index, run identically against both models under the same test conditions. On that index, Opus 5.5 scored 58, ranked first among 212 models evaluated, against GPT-6 Sol's score of 48, ranked eighteenth [10][11]. The same evaluation measured average cost per completed Intelligence Index task at $5.98 for Opus 5.5 against $1.06 for GPT-6 Sol, a reminder that a higher-scoring model evaluated at higher reasoning effort will often cost meaningfully more per task even when its per-token price is not wildly higher [10][11]. GPT-6 Sol also measured faster on raw output speed, 126 tokens per second against Opus 5.5's speed not being published in the same report, though Sol's own time to first token ran considerably higher due to its own reasoning overhead [11].

That is one real, independently run comparison, and it is honest to say it is also the only one currently available at this level of rigor. Unite.AI's own head-to-head write-up, drawing on the same Artificial Analysis data, lands on a genuinely useful framing rather than declaring an outright winner: reach for Opus 5.5 on ambiguous, multi-step work where a wrong step is expensive to catch later, migrations, complex investigations, reports that need to reconcile competing sources, and reach for GPT-6 Sol on work with clear inputs and reliable automated checks, transforming already-structured material, drafting from an approved source packet, or making bounded code changes that already have a test suite to validate against [9]. That framing lines up closely with the benchmark pattern from Step 2: Opus 5.5's biggest advantages over the field show up on long-horizon, ambiguous agentic work, not on narrow, well-specified tasks where a cheaper, faster model can validate its own output against a test.

It is also worth flagging a comparison trap that shows up constantly in coverage of same-day launches like this one: benchmark names that look identical across labs often are not measuring the same thing, or the same version of the same thing. A Terminal-Bench score from one lab's launch materials and a differently-versioned Terminal-Bench score from another's are not directly comparable numbers even when the benchmark name matches exactly, since these suites get revised specifically because models keep beating the previous version. Any comparison chart that puts scores from two different benchmark versions in the same column without noting the version number should be treated with real skepticism, the same standard this post applies to every number in it.

Step 8: A Production Routing Pattern Across the Whole Claude Lineup

With Opus 5.5 now sitting between Sonnet 5 and Fable 5.1 on both price and capability, the practical question for most teams shifts from "which model" to "which model for which step of the pipeline." Sonnet 5 remains the fastest and cheapest tier for short, well-specified tasks, classification, extraction, simple formatting, the same role it played before this release [3]. Opus 5.5 is now the sensible default for the great majority of agentic coding, knowledge work, and computer-use tasks that previously would have needed Opus 5 or, for cost reasons, would have been routed to Fable 5.1 reluctantly. Fable 5.1 remains the right call specifically for the narrow slice of tasks where Opus 5.5's benchmark scores in Step 2 still trail it, most visibly Terminal-Bench-Science and other genuinely open-ended research-style work, or where your own evaluation set shows Opus 5.5 falling short after actually testing it.

python
/code def route_task(task_type: str, is_long_horizon: bool, opus_passed_eval: bool = True) -> dict: """A three-tier router across Sonnet 5, Opus 5.5, and Fable 5.1. Opus 5.5 is now the default for most agentic and knowledge work. Fable 5.1 is reserved for the narrow slice of tasks where your own evaluation shows Opus 5.5 falling short, not a default tier.""" if task_type in ("classification", "extraction", "formatting", "short_qa"): return {"model": "claude-sonnet-5", "effort": "low"} if not opus_passed_eval and is_long_horizon: # Escalate only after Opus 5.5 has actually been tried and measured # against your own evaluation set, per Anthropic's own guidance. return {"model": "claude-fable-5-1", "effort": "high"} effort = "high" if is_long_horizon else "medium" return {"model": "claude-opus-5-5", "effort": effort} print(route_task("short_qa", is_long_horizon=False)) print(route_task("agentic_coding", is_long_horizon=True)) print(route_task("scientific_research", is_long_horizon=True, opus_passed_eval=False))

A few production notes on top of that simplified router. Log the model and effort level used per request, not just the request type, since effort-routing assumptions drift as prompt templates and task distributions change over time. Re-benchmark against your own evaluation set after this migration specifically, since the relative gap between Opus 5.5, Fable 5.1, and Sonnet 5 changed meaningfully from the Opus 5 generation, and a router tuned for the old gap is not guaranteed to route optimally for the new one. Treat Fable 5.1 as an escalation path triggered by a failed validation check on Opus 5.5's output, not a default for an entire task category, since the cost gap between the two tiers is large enough that misrouted traffic adds up fast across real volume.

If you want to explain this benchmark story to a team or an audience as a short generated clip instead of a static table, here is a video generation prompt built around the same mountain-camp altitude metaphor used in the hero image above, written for a Wan or Veo-style video model.

A wide cutaway view of a snow-capped mountain rendered as a clean flat-color technical diagram, camera slowly pushing in from base camp toward the summit. Five small triangular flags are planted at different labeled altitude bands along the slope. As the camera rises past each flag in turn, a soft glowing outline briefly highlights it and a thin dotted line draws itself from the flag to a small altitude scale-bar running down the left edge, revealing its short label and percentage one at a time in order from lowest to highest camp. When the camera reaches the highest flag near the summit, its outline glows brightest and holds steady while the four lower flags dim slightly, and a small price-tag icon fades in beside the highest flag showing a lower dollar figure than the flags below it. Clean vector motion-graphics style throughout, flat muted blues and greys, no photographic lighting, no watermark, no logos, no people, steady smooth camera motion with no shake.

Common Mistakes and Misunderstandings

Assuming Opus 5.5 wins every benchmark. It does not. GPT-6 Astra beats it on AutomationBench by 1.4 points and on Terminal-Bench-Science by nearly 6 points. Both gaps are real and worth knowing if your workload happens to be workflow automation or open-ended scientific research specifically.

Treating the 40 percent cost savings figure as universal. That number describes a typical workload mixing fresh input, cached reads, and generated output. A workload with very little cache reuse, single-shot requests with no repeated system prompt, will see a smaller savings than a long, cache-heavy agentic session, since most of the 40 percent comes from the cache-read cut compounding across many turns, as Step 3's calculator shows directly.

Confusing Fast mode with a higher effort level. Fast mode, at roughly double the standard rate, optimizes for lower latency specifically, not deeper reasoning. Pairing Fast mode with a low effort setting on a genuinely hard task will not reproduce Opus 5.5's benchmark-leading scores, which come from standard mode at appropriate effort levels for the task.

Comparing benchmark scores across different benchmark versions. As Step 7 covers, a Terminal-Bench or OSWorld score measured on one version of the suite is not directly comparable to a score on a different version, even with an identical benchmark name. This trap shows up constantly in third-party summary articles that copy numbers from multiple sources without checking version alignment.

Assuming forced tool use still works after migrating from Opus 5. This is one of the four breaking changes from Step 1, and it fails loudly with an error rather than degrading silently, but only if you actually test the migration path before shipping it, not after.

Ignoring the AutomationBench and Terminal-Bench-Science results because they complicate a clean marketing narrative. A production decision based only on the benchmarks where Opus 5.5 wins, while ignoring the two where it does not, risks routing exactly the wrong workload, open-ended scientific research or complex multi-app workflow automation, to the wrong model.

Production Best Practices for Teams Adopting Opus 5.5

Build your own evaluation set before migrating, not after. General benchmarks like the ones in Step 2 are useful for comparing models in the abstract, but they are not a substitute for testing against your actual task distribution, the same discipline that mattered for every prior Claude generation migration.

Measure your own cache hit rate before trusting the 40 percent savings claim. Log input tokens, output tokens, cache write tokens, and cache read tokens per request separately, not just total spend, so you can see which lever is actually driving your savings and which workloads are not benefiting from the cache-read cut at all.

Set a hard per-session token ceiling regardless of effort level. Even with genuine efficiency gains, a single malformed or ambiguous prompt at a high effort setting can still consume an outsized share of a budget before monitoring catches it.

Treat the classifier screening pipeline as a real operational dependency, not background noise. A production agent with tool access should have a defined fallback behavior for a legitimate request that trips a safety classifier anyway, since "85 percent fewer" attempted boundary circumventions is a real improvement, not a guarantee of zero false positives.

Re-test any workflow that programmatically edits earlier conversation turns. Given that editing prior turns can silently invalidate preserved thinking blocks, any eval harness or agent framework that prunes or rewrites context needs a specific regression test for this exact scenario before it ships against Opus 5.5 in production.

Where This Fits Into a Real Content Pipeline

None of this is unique to software engineering teams. Any multi-step generation pipeline has the identical shape of decision baked into it: which steps genuinely need the strongest available reasoning, and which steps were already solved by a cheaper, faster tier months ago. A content pipeline that turns a topic into a finished video, the kind of workflow behind Text2Shorts in Miraflow AI, is a real example of exactly this pattern. Script generation from a raw topic benefits from stronger reasoning about pacing and structure, the same category of long-horizon, ambiguous task where Opus 5.5's benchmark advantage over cheaper tiers is largest. Turning a finished script into scene-by-scene visual prompts is closer to a well-specified formatting task, where a cheaper, faster model captures most of the value at a fraction of the cost. The same tiering logic applies to the AI Image Generator in Miraflow AI and the cinematic AI video generator, where a planning step and a mechanical execution step have meaningfully different reasoning requirements even though both feed into the same finished piece of content. If you want the deeper mechanical explanation of how effort-based routing works underneath a model like this, our breakdown of Claude Opus 5 vs Sonnet 5 covers the effort toggle in detail, and our Claude Fable 5.1 explainer covers where the top of this lineup sits when a task genuinely needs it. You can browse more explainers like this on the Miraflow AI blog, and the same reasoning-versus-speed tradeoff shows up in how the AI Music Generator on Miraflow AI balances its own generation quality against turnaround time.

Frequently Asked Questions

Is Claude Opus 5.5 always better than Claude Fable 5.1? No. Opus 5.5 matches or beats Fable 5.1 on most published benchmarks while costing roughly 60 percent less per token, but Fable 5.1 remains Anthropic's stronger model on the small set of tasks where the benchmark gap still favors it, and Anthropic's own guidance is to default to the cheaper model and escalate only where your own evaluation shows it falling short.

Does Opus 5.5 really beat GPT-6 Astra on most benchmarks? On the benchmarks Anthropic published, yes on most, including Terminal-Bench 4.0, FrontierCode v1.1, GDPval-AA v2.1, and Humanity's Last Exam. GPT-6 Astra wins on AutomationBench and by a wider margin on Terminal-Bench-Science, so "most" is accurate, "all" is not.

Is Claude Opus 5.5 cheaper than GPT-6 Sol? No, not on raw list price. GPT-6 Sol runs $2 input and $10 output per million tokens against Opus 5.5's $4 and $20, undercutting it on both. The one available independent comparison, Artificial Analysis's Intelligence Index, shows Opus 5.5 scoring higher at a higher average cost per task, so the right choice depends on whether the task needs the extra reasoning quality or has cheap, reliable ways to check GPT-6 Sol's output instead.

Can I disable thinking on Opus 5.5 to save cost like I could on some earlier models? No. Adaptive thinking is always on for Opus 5.5, one of the model's four breaking changes from the Opus 5 generation. The effort parameter, not a thinking on/off toggle, is now the only lever for controlling reasoning depth and cost.

Is Opus 5.5 available on AWS Bedrock and the Claude Platform on AWS? Yes. It is available on both as anthropic.claude-opus-5-5 on Bedrock and claude-opus-5-5 on the Claude Platform on AWS, alongside Google Cloud and Microsoft Azure, from launch day.

Should a multi-step content pipeline route every step through Opus 5.5? Generally no. Long-horizon, structurally demanding steps like script or narrative generation are where Opus 5.5's advantage over a cheaper tier like Sonnet 5 is largest. Well-specified formatting, extraction, or metadata steps show a much narrower advantage, and routing those steps through the most capable available model anyway is one of the most common ways teams overspend after any model upgrade.

Conclusion

Claude Opus 5.5 is a genuinely unusual kind of release: a cheaper model in a lab's own lineup that beats the more expensive one on most, though honestly not all, of the benchmarks that matter for real agentic work, while costing roughly 60 percent less per token than that pricier sibling and about 40 percent less than the model it directly replaces. The mechanism behind that is not just a price cut, it is a real token-efficiency improvement, visible in the 2.5x token reduction on the 200,000-line audit case study and the concrete HAProxy migration comparison against Fable 5.1. That it shipped the same day OpenAI cut its own prices by half with GPT-6 Sol and Luna makes the moment sharper, but it does not make either release a clean winner over the other, since the two labs are optimizing for genuinely different points on the cost-versus-reasoning curve. The teams that get real value out of this release will be the ones who test it against their own workload, route deliberately between Sonnet 5, Opus 5.5, and Fable 5.1 based on real evaluation results rather than sticker price alone, and treat the two benchmarks where GPT-6 Astra still leads as useful signal rather than an inconvenient footnote to ignore.

References and Sources

[1] Anthropic. "Introducing Claude Opus 5.5."

[2] Anthropic. "Claude Opus 5.5."

[3] Anthropic Platform Docs. "Claude Opus 5.5 Overview."

[4] Anthropic Platform Docs. "What's New in Claude Opus 5.5."

[5] AWS Machine Learning Blog. "Claude Opus 5.5 is now available on AWS."

[6] TechCrunch. "Anthropic releases Opus 5.5 with lower prices and Fable-level performance."

[7] VentureBeat. "Anthropic releases Claude Opus 5.5, beating Fable 5.1 on key agentic benchmarks at 60% cheaper API price."

[8] Unite.AI. "Anthropic Releases Claude Opus 5.5 With Lower Pricing and New Safeguards."

[9] Unite.AI. "Claude Opus 5.5 vs GPT-6 Sol: Which Is Better for Your Business?"

[10] Artificial Analysis. "Claude Opus 5.5: Intelligence, Performance & Price Analysis."

[11] Artificial Analysis. "GPT-6 Sol: Intelligence, Performance & Price Analysis."

[12] VentureBeat. "OpenAI releases GPT-6 Sol and Luna models, slashing API costs 50% or more."

[13] OpenAI. "Introducing GPT-6 Sol and Luna."

[14] MacRumors. "Anthropic Launches Claude Opus 5.5 With Fable-Level Performance at a Lower Price."

[15] SiliconANGLE. "Anthropic releases Claude Opus 5.5 and OpenAI counters with two cheaper GPT-6 models."

[16] 9to5Mac. "Anthropic upgrades Claude with new Opus 5.5 model, details here."

[17] MarkTechPost. "Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5."

[18] Trending Topics. "Claude Opus 5.5: Anthropic Launches New Top Model Despite Calling for AI Slowdown."

[19] Anthropic. "Claude Opus 5.5 System Card."

[20] Anthropic Platform Docs. "Fast Mode."

[21] BNN Bloomberg. "Anthropic unveils Claude Opus 5.5."