Everyone Quoted the Benchmark. The Price Was the Announcement.

Anthropic launched Claude Opus 5.5 on September 22, 2026. It leads Terminal-Bench 4.0 by 8.5 points, which is a fine headline and not the important number. Input and output tokens dropped 20%, cache reads dropped 60%, and typical workloads cost about 40% less than they did on Opus 5. That is the part that changes what you are allowed to build.

What Anthropic Actually Shipped

On September 22, 2026, Anthropic released Claude Opus 5.5, the first model in a new 5.5 family and its current flagship. It shipped simultaneously on the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry. The API identifier is claude-opus-5-5, the knowledge cutoff is June 2026, the context window goes up to 1M tokens, and maximum output is 128K tokens — 300K on the Batch API behind a beta header.

The coverage that followed split into two camps, and both missed the same thing. One camp led with Terminal-Bench. The other led with "performs at Fable 5.1 level for less money," which is closer but still frames price as a footnote to capability.

Read the pricing table instead. It is the product decision.

−20% Input and output tokens: $4 and $20 per million, down from $5 and $25
−60% Cache reads: $0.20 per million, down from $0.50
−40% Combined effect on a typical workload at default settings
+30% Output generation speed compared with Opus 5

The Full Price Table

Rate (per 1M tokens)Opus 5Opus 5.5Change
Input$5.00$4.00−20%
Output$25.00$20.00−20%
Cache read$0.50$0.20−60%
Cache write (5 min)$6.25$5.00−20%
Cache write (1 hour)$10.00$8.00−20%
Batch API$2.50 / $12.50$2.00 / $10.00−20%
Fast mode—$8.00 / $40.00New, research preview

Every line moved 20% except one. Cache reads moved 60%, and that asymmetry is deliberate.

Why the Cache Read Line Is the Whole Story

If you have never costed an agent, the intuition is that you pay mostly for what the model writes. You do not. You pay for what it re-reads.

An agent loop works like this. Turn one sends a system prompt, a tool catalogue, relevant files and the user's request. The model calls a tool. Turn two sends all of that again, plus the tool result. Turn three sends everything again, plus the next result. A twenty-step task re-transmits the same context twenty times, and prompt caching exists precisely so those repeats bill at the cache read rate instead of the full input rate.

So the cost of a long-running agent is dominated by a number most pricing announcements bury in a sub-row. Anthropic cut that number by 60% while cutting everything else by 20%. That is not a uniform discount. That is a targeted subsidy on the one workload shape the company is betting on.

The shorthand. A 20% cut is a discount on chatbots. A 60% cut on cache reads is a discount on agents. The headline number Anthropic quotes — roughly 40% cheaper on typical workloads — is what you get when those two compound on a real agentic trace.

The Benchmarks, With the Caveat Attached

Opus 5.5 posts a top-three result across most public evaluations, and a clear lead on agentic coding specifically. It is not a sweep.

BenchmarkOpus 5.5Best competitorResult
Terminal-Bench 4.066.4%57.9% — GPT-6 AstraLead, +8.5
CursorBench 4.057.8%51.8% — Fable 5.1Lead, +6.0
GDPval-AA v2.11846 Elo1735 — Fable 5.1Lead, +111
FrontierCode v1.154.4%53.3% — GPT-6 AstraLead, +1.1
OSWorld 2.081.8%80.7% — Fable 5.1Lead, +1.1
AutomationBench40.0%41.4% — GPT-6 AstraBehind, −1.4
Terminal-Bench-Science 0.158.7%64.6% — GPT-6 AstraBehind, −5.9

Two things are worth saying plainly. First, most of these scores are reported at maximum effort, which is not the setting the 40% cost figure refers to — that one is measured at default effort. You do not get the benchmark numbers and the cost numbers in the same configuration. Second, four of the seven leads are inside six points, and two are inside 1.2 points. On benchmarks with this much run-to-run variance, that is a tie dressed up as a win.

The Terminal-Bench 4.0 gap is the exception and the one that matters: 8.5 points on long-horizon terminal work, which is exactly the task shape an autonomous coding agent runs all day.

01
The Economics Read

Capability Is Converging. Price Is Where the Competition Moved.

Count what shipped in three weeks: Claude Fable 5.1 and Mythos 5.1 on September 1, GPT-6 Astra on September 3, then Gemini 3.8 Flash, DeepSeek V4.1-Flash, and now Opus 5.5 on September 22. Five frontier launches, and the spread between the best and third-best model on most benchmarks is now smaller than the run-to-run noise. When capability converges, the differentiator stops being the score and becomes cost per unit of useful work. That is what a commodity market looks like from the inside, and it is a very good thing for buyers: the labs are now competing on the axis you actually feel on an invoice. It also means "which model is smartest" has quietly become the wrong procurement question. The right one is which model finishes your specific task at the lowest total cost, and that answer is increasingly workload-specific rather than leaderboard-specific.

What a 40% Cut Actually Unblocks

Every engineering organization has a drawer of agent ideas that died in a spreadsheet rather than in a design review. The pattern is always the same: the thing worked, and then somebody divided total tokens by tickets handled.

None of those use cases changed this week. Nothing about the requirement, the design or the risk moved. Only the denominator did — and for the cache-heavy ones, it moved by more than the headline 40%.

That recalculation is happening right now, quietly, inside finance decks rather than on engineering blogs. It is the least visible and most consequential effect of this launch.

Five Breaking Changes Nobody Put in the Headline

Opus 5.5 is not a drop-in replacement for Opus 5. If you are migrating, these surface as production errors rather than as degraded output:

  1. Thinking cannot be disabled. It is adaptive by default. You either use adaptive thinking or omit the parameter; explicitly turning it off is no longer an option.
  2. Forced tool use is gone. tool_choice: "any" now returns a 400. Any orchestration layer that relied on guaranteeing a tool call needs restructuring.
  3. Thinking blocks are bound to their conversation. Replaying them after the system prompt has changed returns a 400. Caching and replay layers that rewrite system prompts between turns will break.
  4. The older computer-use toolset is rejected. You must move to computer_toolset_20260801.
  5. Internal notes now surface as thinking blocks — empty at the default display setting. This one is silent, which makes it the one that costs an afternoon.

On the subscription side, five-hour session limits rose about 20% on Pro, Max, Team and seat-based Enterprise plans, and subscribers received a one-off rate-limit reset usable through October 22. Anthropic's earliest retirement commitment for the model is September 22, 2027 or later.

What is still unverified. The 40% figure is Anthropic's own, measured on its own definition of a typical workload at default effort, and no independent reproduction existed at the time of writing. Benchmark numbers are vendor-published and mostly run at maximum effort. Fast mode is a research preview limited to the Claude API. Price your own trace before committing a budget to any of these numbers — caching hit rates vary enormously between workloads, and a low hit rate erases most of the advertised gain.

The Part Nobody Is Pricing In

Cost per token is not only a line item. It is an attack surface parameter.

Every piece of adversarial AI research this year has assumed agents running continuously and at scale. The PaperCut swarm that breached 395 organizations ran many agents in parallel over an extended period. Midnight Blizzard's malware rewriting loop depended on repeated regeneration against a detection signal. The UK AISI work on self-assembled attack chains is fundamentally about long-horizon autonomous operation. All of it has a compute bill.

The same 60% cut on cache reads that makes your internal agent viable makes a long-running adversarial agent viable too. Continuous reconnaissance against a target — enumerating, probing, re-reading the same accumulated context on every iteration — is precisely the cache-heavy shape that just got cheapest. And the attacker is not paying for your compliance review, your evaluation harness or your staged rollout.

This is not an argument against cheaper inference. It is an argument that defensive budgeting has to track the same curve. If your threat model assumed automated attacks were rate-limited by cost, that assumption depreciated 40% in one release, and it will depreciate again at the next one.

02
The Security Read

Defenders Buy Inference on a Budget Cycle. Attackers Buy It on a Credit Card.

There is a structural asymmetry in how the two sides absorb a price cut. An enterprise that wants to deploy a newly-affordable agent has to get it through architecture review, data governance, a security assessment, a procurement cycle and a staged rollout — realistically a quarter, often two. An attacker who wants to run the same workload changes a model string in a config file. Both parties received the same discount on September 22. Only one of them can act on it the same afternoon. That gap is not new, but it widens every time frontier capability gets cheaper, because the constraint that used to bind both sides equally — raw compute cost — is the one being removed. The controls that still bind are the ones that do not depend on the adversary's budget: least privilege, egress restrictions, short-lived credentials, and monitoring the actions an agent takes rather than the tokens it spends.

What To Do With This

1. Reprice the Workloads You Already Rejected

Go back to whatever you shelved on unit economics in the last two quarters and rerun the arithmetic. Weight the cache read line heavily: if the workload re-reads a large stable context across many turns, your real reduction is well above 40%. If it is single-turn with fresh context every time, it is closer to 20%, and the business case may still not close.

2. Measure Your Actual Cache Hit Rate Before Believing Any of This

The advertised saving assumes caching works on your traffic. It often does not. Prompt ordering, system prompt churn between turns and cache TTL expiry all silently drop you back to full input pricing. Instrument the cache-read-to-input token ratio in production before you build a forecast on it.

3. Plan the Migration as a Migration

Five breaking changes, one of them silent. Do not swap the model string in production and hope. The tool_choice: "any" removal in particular will break orchestration code that has worked for two model generations, and the thinking-block binding rule breaks context-management layers that rewrite system prompts mid-conversation.

4. Decide Where Effort Actually Belongs

Effort is now the main cost lever, defaulting to medium. Benchmarks are quoted at maximum, savings at default. Treat effort as a per-route setting, not a global one: maximum on the small number of hard tasks that justify it, default or below everywhere else. Getting this wrong in either direction is the easiest way to lose the entire price cut.

5. Re-Baseline the Threat Model, Not Just the Budget

If any of your risk documentation says or implies that a class of automated attack is impractical at scale, check whether "at scale" was a cost argument. Most of them are. Then apply the same reduction and ask whether the conclusion survives.

6. Add Agent Spend to Your Inventory, Alongside Agent Access

Cheaper agents means more agents, deployed faster, by more teams, with less scrutiny. That is the real second-order risk of this launch — not the model, but the volume of new autonomous software about to enter production on a business case that finally closed. Know what is running, what it can reach, and who approved it. The plugin and supply chain failures documented in Plugin4Shell and GitSpawn all assumed an install base. The install base is about to grow.

The interesting thing about Opus 5.5 is not that it is smarter. It is that "smarter" stopped being the axis. Five frontier models launched in three weeks, clustered within noise of each other on most evaluations, and the one that led did so by cutting the price of the specific token type that agent workloads consume most. Capability is becoming table stakes. Cost structure is becoming the product.

Frequently Asked Questions

How much does Claude Opus 5.5 cost?
$4 per million input tokens and $20 per million output tokens, down from $5 and $25 on Opus 5. Cache reads are $0.20 per million, down from $0.50. Cache writes are $5 for the 5-minute TTL and $8 for the 1-hour TTL. The Batch API is $2 input and $10 output. A separate fast mode, currently a research preview on the Claude API only, costs $8 and $40.
Where does the 40% cost reduction come from?
It is a blended figure, not a single line item. Input and output fell 20%, cache reads fell 60%, and Anthropic reports roughly 40% lower cost on typical workloads at default effort once those compound. How close you get to 40% depends almost entirely on your cache hit rate: cache-heavy agentic traces can exceed it, while single-turn requests with fresh context every time see closer to the 20% headline cut.
Is Opus 5.5 actually better than GPT-6 Astra?
On agentic coding, yes and by a clear margin: 66.4% against 57.9% on Terminal-Bench 4.0. Elsewhere it is mixed. Opus 5.5 leads FrontierCode v1.1 and OSWorld 2.0 by roughly a point, which is inside benchmark noise, and trails GPT-6 Astra on AutomationBench (40.0% vs 41.4%) and Terminal-Bench-Science 0.1 (58.7% vs 64.6%). Most published scores are at maximum effort, which is not the configuration the cost figures describe.
What breaks if I migrate from Opus 5?
Five things. Thinking can no longer be disabled and is adaptive by default. tool_choice: "any" returns a 400, so forced tool use is gone. Replaying thinking blocks after the system prompt has changed returns a 400. The older computer-use toolset is rejected in favour of computer_toolset_20260801. And internal notes now appear as thinking blocks, empty at the default display setting, which is the change most likely to surprise you silently.
What is the context window and output limit?
Up to 1M tokens of context, with a maximum output of 128K tokens, extending to 300K on the Batch API behind a beta header. The knowledge cutoff is June 2026. The API identifier is claude-opus-5-5, and it is available on the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry.
Did subscription limits change?
Yes. Five-hour session limits rose roughly 20% on Pro, Max, Team and seat-based Enterprise plans as of September 22, 2026, and subscribers received a one-off rate-limit reset usable through October 22. Anthropic's earliest retirement commitment for Opus 5.5 is September 22, 2027 or later.
Why does a price cut matter for security?
Because most arguments that an automated attack is impractical at scale are cost arguments underneath. Long-running adversarial workloads such as continuous reconnaissance are cache-heavy by nature, so they benefit disproportionately from the 60% cache read reduction. Defenders absorb a price cut through procurement and rollout cycles measured in quarters; an attacker absorbs it by editing a config file. The controls that still hold are the ones independent of the adversary's budget: least privilege, egress control, short-lived credentials and action-level monitoring.

My Take

I spent most of this year writing about what goes wrong around models rather than inside them — an unverified commit pin in four coding agents, a swarm breaching 395 organizations, agents building a communication channel out of a package repository. This launch is not one of those stories, and that is exactly why it is worth reading carefully.

Opus 5.5 is a good model. It is also, more importantly, a signal about where the market went. Five frontier launches in three weeks, benchmark spreads inside noise, and the winning move was a targeted cut on the token type agents consume most. Nobody at any of these labs believes the next point of benchmark is what closes an enterprise deal. The pricing page says so more honestly than the model card does.

The thing I would push back on is the reflex to treat that as unambiguously good news. Cheaper inference is good news for the people who were already going to do this carefully. It is neutral-to-bad news for everyone who was being restrained by a budget rather than by a policy, and in most organizations I have seen, the budget was doing more governance work than anyone wants to admit. When the constraint that was accidentally enforcing discipline gets removed, the discipline does not appear on its own.

So the practical version of this article is short. Rerun your numbers, because they genuinely changed. Instrument your cache hit rate, because the advertised saving is conditional on something most teams never measure. Plan the migration properly, because five breaking changes is not a version bump. And before you green-light the wave of agents that just became affordable, write down what each of them can reach.

The cost of running an agent fell 40% this week. The cost of an agent doing something you did not intend did not move at all.

Related Articles:

Kodjo Apedoh

About the Author

Kodjo Apedoh

Network Engineer & AI Entrepreneur

Founder of TechVernia & SankaraShield. Certified Network Security Engineer with 4+ years of experience specializing in network automation (Python), AI tools research, and advanced security implementations. Also builds iOS and Android applications. Holds certifications from Palo Alto Networks, Fortinet, and Cisco. Based in Arlington, Virginia.

Connect on LinkedIn →